Ross ROSS = Recommend OSS · open-source software intelligence for agents

hendrycks/test resource

Measuring Massive Multitask Language Understanding | ICLR 2021 observed · 2026-08-28

github.com/hendrycks/test · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2186
  • days_rel: n/a
  • days_push: 1193
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1613 stars · 119 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

The MMLU (Massive Multitask Language Understanding) benchmark dataset and evaluation code from Hendrycks et al., ICLR 2021. It contains multiple-choice questions spanning 57 academic and professional subjects, plus OpenAI API evaluation scripts for measuring language model accuracy.

Use cases

  • evaluate an LLM on MMLU benchmark
  • measure multitask knowledge of a language model
  • compare my model's accuracy against GPT-3 leaderboard scores
  • run few-shot evaluation of a text model on 57 subjects
  • download MMLU test questions for model evaluation

When to choose

  • you need a standardized academic knowledge benchmark for LLM evaluation
  • you want to compare your model against published MMLU leaderboard results
  • you need multiple-choice evaluation data across humanities, STEM, and social sciences

When to avoid

  • you need a live leaderboard or modern harness — the repo's evaluation code is dated and most teams now use lm-evaluation-harness
  • you need training data rather than an evaluation test set
  • you need benchmarks beyond multiple-choice academic questions

Facets

dataset · maturity maintenance

machine-learning benchmarking nlp large-language-models machine-learning artificial-intelligence education python mmlu llm-evaluation few-shot-learning benchmark-dataset gpt-3 multiple-choice

6 sources

Member repositories

RepositoryRoleHealth v2
hendrycks/testmain32

For agents

markdown · JSON · MCP: product_card(name="hendrycks/test")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem