# FranxYao/chain-of-thought-hub

Benchmarking large language models' complex reasoning ability with chain-of-thought prompting

Repository: https://github.com/FranxYao/chain-of-thought-hub
Canonical: https://ross.abutalabs.com/products/chain-of-thought-hub
Language: Jupyter Notebook
License: MIT
License Family: permissive
Last push: 2024-08-04T09:40:18+00:00

## Health v2 (maintenance only)
Score: 30/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 90
- inputs: {"age_days": 1272, "days_push": 759, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2775, forks 144 (observed 2026-08-28T04:07:21.183247+00:00)

## What it is
A benchmarking hub and collection of evaluation scripts measuring large language models' complex reasoning ability using chain-of-thought prompting across tasks like GSM8K, MATH, BBH, MMLU, and HumanEval. It accompanies an academic paper and tracks model performance leaderboards over time.

## Use cases
- benchmark llm reasoning ability
- compare chain-of-thought prompting performance across models
- evaluate a model on gsm8k and math datasets
- track which llm is best at complex reasoning
- find chain-of-thought prompt examples for evaluation
- measure long-context reasoning of large language models

## When to choose
- you need standardized complex-reasoning benchmarks for LLMs
- you want curated chain-of-thought prompts and eval scripts
- you are comparing open and closed models on math, coding, or knowledge tasks

## When to avoid
- you need a production inference or serving tool
- you want general-purpose LLM evaluation beyond reasoning tasks
- you need actively maintained benchmarks with the newest datasets

## Facets
- artifact type: dataset
- maturity: maintenance
- function: benchmarking, prompt-engineering, machine-learning
- domain: large-language-models, artificial-intelligence, tutorials
- platform: python
- tags: chain-of-thought, llm-evaluation, reasoning-benchmarks, gsm8k, math-benchmark, jupyter-notebooks, natural-language-processing

## Member repositories
- FranxYao/chain-of-thought-hub (main) score 30

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:21.183247+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T08:16:33.162618+00:00, confidence not recorded.
  - readme: https://github.com/FranxYao/chain-of-thought-hub (fetched 2026-08-28T04:07:21.183247+00:00, sha 339fe4856af6)
- Data as of 2026-08-30T08:39:29.467469+00:00.
