openai/simple-evals
None observed · 2026-08-28
Health v2 · maintenance only
60/100
- Activity 78
- Release rhythm 35
- Longevity 62
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 874
- days_rel: n/a
- days_push: 133
- n_releases_24m: 0
Adoption not part of the score
4612 stars · 506 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
A lightweight Python library from OpenAI for evaluating language models against benchmarks like MMLU, GPQA, MATH, HumanEval, SimpleQA, HealthBench, and BrowseComp. It hosts reference implementations used to transparently publish model accuracy numbers.
Use cases
- evaluate an LLM on MMLU or GPQA benchmarks
- measure model factual accuracy with SimpleQA
- run HealthBench evaluations on a language model
- reproduce OpenAI's published benchmark numbers
- compare model performance on math and coding tasks
- benchmark a new model against o3 or o4-mini results
When to choose
- you want lightweight, reference-quality LLM evaluation code
- you need to reproduce OpenAI's published benchmark scores
- you want to evaluate models on SimpleQA, HealthBench, or BrowseComp
When to avoid
- you need a full-featured, actively maintained eval framework with new benchmarks
- you need evaluations updated for the latest models, since the repo is deprecated for new results
- you need a production evaluation pipeline with extensive integrations
Facets
library · maturity maintenance
machine-learning benchmarking llm-inference large-language-models machine-learning developer-tools python llm-evaluation benchmarks openai simpleqa healthbench browsecomp
1 source
- readme: https://github.com/openai/simple-evals · fetched 2026-08-28 · e97a732b51b2
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| openai/simple-evals | main | 60 |
For agents
markdown · JSON · MCP: product_card(name="openai/simple-evals")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem