Ross ROSS = Recommend OSS · open-source software intelligence for agents

FranxYao/chain-of-thought-hub resource

Benchmarking large language models' complex reasoning ability with chain-of-thought prompting observed · 2026-08-28

github.com/FranxYao/chain-of-thought-hub · Jupyter Notebook · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

30/100

  • Activity 0
  • Release rhythm 35
  • Longevity 90

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1272
  • days_rel: n/a
  • days_push: 759
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2775 stars · 144 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A benchmarking hub and collection of evaluation scripts measuring large language models' complex reasoning ability using chain-of-thought prompting across tasks like GSM8K, MATH, BBH, MMLU, and HumanEval. It accompanies an academic paper and tracks model performance leaderboards over time.

Use cases

  • benchmark llm reasoning ability
  • compare chain-of-thought prompting performance across models
  • evaluate a model on gsm8k and math datasets
  • track which llm is best at complex reasoning
  • find chain-of-thought prompt examples for evaluation
  • measure long-context reasoning of large language models

When to choose

  • you need standardized complex-reasoning benchmarks for LLMs
  • you want curated chain-of-thought prompts and eval scripts
  • you are comparing open and closed models on math, coding, or knowledge tasks

When to avoid

  • you need a production inference or serving tool
  • you want general-purpose LLM evaluation beyond reasoning tasks
  • you need actively maintained benchmarks with the newest datasets

Facets

dataset · maturity maintenance

benchmarking prompt-engineering machine-learning large-language-models artificial-intelligence tutorials python chain-of-thought llm-evaluation reasoning-benchmarks gsm8k math-benchmark jupyter-notebooks natural-language-processing

1 source

Member repositories

RepositoryRoleHealth v2
FranxYao/chain-of-thought-hubmain30

For agents

markdown · JSON · MCP: product_card(name="FranxYao/chain-of-thought-hub")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem