FranxYao/chain-of-thought-hub resource
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting observed · 2026-08-28
Health v2 · maintenance only
30/100
- Activity 0
- Release rhythm 35
- Longevity 90
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1272
- days_rel: n/a
- days_push: 759
- n_releases_24m: 0
Adoption not part of the score
2775 stars · 144 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A benchmarking hub and collection of evaluation scripts measuring large language models' complex reasoning ability using chain-of-thought prompting across tasks like GSM8K, MATH, BBH, MMLU, and HumanEval. It accompanies an academic paper and tracks model performance leaderboards over time.
Use cases
- benchmark llm reasoning ability
- compare chain-of-thought prompting performance across models
- evaluate a model on gsm8k and math datasets
- track which llm is best at complex reasoning
- find chain-of-thought prompt examples for evaluation
- measure long-context reasoning of large language models
When to choose
- you need standardized complex-reasoning benchmarks for LLMs
- you want curated chain-of-thought prompts and eval scripts
- you are comparing open and closed models on math, coding, or knowledge tasks
When to avoid
- you need a production inference or serving tool
- you want general-purpose LLM evaluation beyond reasoning tasks
- you need actively maintained benchmarks with the newest datasets
Facets
dataset · maturity maintenance
benchmarking prompt-engineering machine-learning large-language-models artificial-intelligence tutorials python chain-of-thought llm-evaluation reasoning-benchmarks gsm8k math-benchmark jupyter-notebooks natural-language-processing
1 source
- readme: https://github.com/FranxYao/chain-of-thought-hub · fetched 2026-08-28 · 339fe4856af6
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| FranxYao/chain-of-thought-hub | main | 30 |
For agents
markdown · JSON · MCP: product_card(name="FranxYao/chain-of-thought-hub")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem