Ross ROSS = Recommend OSS · open-source software intelligence for agents

stanford-crfm/helm

Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models. observed · 2026-08-28

github.com/stanford-crfm/helm · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

87/100

  • Activity 95
  • Release rhythm 69
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 32
  • age_days: 1738
  • days_rel: 125
  • days_push: 33
  • n_releases_24m: 14

Full methodology

Adoption not part of the score

2887 stars · 408 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

HELM is a Python framework from Stanford's CRFM for holistic, reproducible evaluation of foundation models including LLMs and multimodal models. It provides standardized benchmarks, a unified model interface across providers, metrics beyond accuracy, and a web UI/leaderboard for inspecting and comparing results.

Use cases

  • evaluate LLMs on benchmarks like MMLU-Pro and GPQA
  • compare models across providers on a leaderboard
  • measure model bias, toxicity, and efficiency beyond accuracy
  • inspect individual prompts and model responses in a web UI
  • run reproducible model evaluation suites
  • benchmark multimodal foundation models

When to choose

  • you need rigorous, reproducible benchmarking of LLMs or multimodal models
  • you want standardized datasets and metrics maintained by a research lab
  • you need a unified interface to evaluate models from OpenAI, Anthropic, Google, and others
  • you want a web leaderboard and UI for exploring evaluation results

When to avoid

  • you need actively developed new features (HELM entered maintenance mode in June 2026)
  • you only need lightweight eval harnesses for a single model
  • you need non-Python tooling or real-time production monitoring

Facets

framework · maturity maintenance

benchmarking machine-learning llm-inference data-science testing large-language-models machine-learning artificial-intelligence developer-tools analytics python cli cross-platform llm-evaluation foundation-models leaderboard benchmarks multimodal-evaluation reproducibility web-server

2 sources

Member repositories

RepositoryRoleHealth v2
stanford-crfm/helmmain87

For agents

markdown · JSON · MCP: product_card(name="stanford-crfm/helm")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem