stanford-crfm/helm
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models. observed · 2026-08-28
Health v2 · maintenance only
87/100
- Activity 95
- Release rhythm 69
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 32
- age_days: 1738
- days_rel: 125
- days_push: 33
- n_releases_24m: 14
Adoption not part of the score
2887 stars · 408 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
HELM is a Python framework from Stanford's CRFM for holistic, reproducible evaluation of foundation models including LLMs and multimodal models. It provides standardized benchmarks, a unified model interface across providers, metrics beyond accuracy, and a web UI/leaderboard for inspecting and comparing results.
Use cases
- evaluate LLMs on benchmarks like MMLU-Pro and GPQA
- compare models across providers on a leaderboard
- measure model bias, toxicity, and efficiency beyond accuracy
- inspect individual prompts and model responses in a web UI
- run reproducible model evaluation suites
- benchmark multimodal foundation models
When to choose
- you need rigorous, reproducible benchmarking of LLMs or multimodal models
- you want standardized datasets and metrics maintained by a research lab
- you need a unified interface to evaluate models from OpenAI, Anthropic, Google, and others
- you want a web leaderboard and UI for exploring evaluation results
When to avoid
- you need actively developed new features (HELM entered maintenance mode in June 2026)
- you only need lightweight eval harnesses for a single model
- you need non-Python tooling or real-time production monitoring
Facets
framework · maturity maintenance
benchmarking machine-learning llm-inference data-science testing large-language-models machine-learning artificial-intelligence developer-tools analytics python cli cross-platform llm-evaluation foundation-models leaderboard benchmarks multimodal-evaluation reproducibility web-server
2 sources
- readme: https://github.com/stanford-crfm/helm · fetched 2026-08-28 · 1740901d80cf
- homepage: https://crfm.stanford.edu/helm · fetched 2026-08-29 · 8bab84344c16
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| stanford-crfm/helm | main | 87 |
For agents
markdown · JSON · MCP: product_card(name="stanford-crfm/helm")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem