Ross ROSS = Recommend OSS · open-source software intelligence for agents

openai/simple-evals

None observed · 2026-08-28

github.com/openai/simple-evals · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

60/100

  • Activity 78
  • Release rhythm 35
  • Longevity 62

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 874
  • days_rel: n/a
  • days_push: 133
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

4612 stars · 506 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A lightweight Python library from OpenAI for evaluating language models against benchmarks like MMLU, GPQA, MATH, HumanEval, SimpleQA, HealthBench, and BrowseComp. It hosts reference implementations used to transparently publish model accuracy numbers.

Use cases

  • evaluate an LLM on MMLU or GPQA benchmarks
  • measure model factual accuracy with SimpleQA
  • run HealthBench evaluations on a language model
  • reproduce OpenAI's published benchmark numbers
  • compare model performance on math and coding tasks
  • benchmark a new model against o3 or o4-mini results

When to choose

  • you want lightweight, reference-quality LLM evaluation code
  • you need to reproduce OpenAI's published benchmark scores
  • you want to evaluate models on SimpleQA, HealthBench, or BrowseComp

When to avoid

  • you need a full-featured, actively maintained eval framework with new benchmarks
  • you need evaluations updated for the latest models, since the repo is deprecated for new results
  • you need a production evaluation pipeline with extensive integrations

Facets

library · maturity maintenance

machine-learning benchmarking llm-inference large-language-models machine-learning developer-tools python llm-evaluation benchmarks openai simpleqa healthbench browsecomp

1 source

Member repositories

RepositoryRoleHealth v2
openai/simple-evalsmain60

For agents

markdown · JSON · MCP: product_card(name="openai/simple-evals")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem