# stanford-crfm/helm

Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.

Repository: https://github.com/stanford-crfm/helm
Canonical: https://ross.abutalabs.com/products/stanford-crfm-helm
Homepage: https://crfm.stanford.edu/helm
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2026-08-01T01:23:17+00:00

## Health v2 (maintenance only)
Score: 87/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 95, release rhythm 69, longevity 100
- inputs: {"age_days": 1738, "days_push": 33, "days_rel": 125, "gap_med": 32, "n_releases_24m": 14}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2887, forks 408 (observed 2026-08-28T04:07:28.251446+00:00)

## What it is
HELM is a Python framework from Stanford's CRFM for holistic, reproducible evaluation of foundation models including LLMs and multimodal models. It provides standardized benchmarks, a unified model interface across providers, metrics beyond accuracy, and a web UI/leaderboard for inspecting and comparing results.

## Use cases
- evaluate LLMs on benchmarks like MMLU-Pro and GPQA
- compare models across providers on a leaderboard
- measure model bias, toxicity, and efficiency beyond accuracy
- inspect individual prompts and model responses in a web UI
- run reproducible model evaluation suites
- benchmark multimodal foundation models

## When to choose
- you need rigorous, reproducible benchmarking of LLMs or multimodal models
- you want standardized datasets and metrics maintained by a research lab
- you need a unified interface to evaluate models from OpenAI, Anthropic, Google, and others
- you want a web leaderboard and UI for exploring evaluation results

## When to avoid
- you need actively developed new features (HELM entered maintenance mode in June 2026)
- you only need lightweight eval harnesses for a single model
- you need non-Python tooling or real-time production monitoring

## Facets
- artifact type: framework
- maturity: maintenance
- function: benchmarking, machine-learning, llm-inference, data-science, testing
- domain: large-language-models, machine-learning, artificial-intelligence, developer-tools, analytics
- platform: python, cli, cross-platform
- tags: llm-evaluation, foundation-models, leaderboard, benchmarks, multimodal-evaluation, reproducibility, web-server

## Member repositories
- stanford-crfm/helm (main) score 87

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:28.251446+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:35:22.532943+00:00, confidence not recorded.
  - readme: https://github.com/stanford-crfm/helm (fetched 2026-08-28T04:07:28.251446+00:00, sha 1740901d80cf)
  - homepage: https://crfm.stanford.edu/helm (fetched 2026-08-29T09:50:36.719675+00:00, sha 8bab84344c16)
- Data as of 2026-08-30T08:39:29.467469+00:00.
