EleutherAI/lm-evaluation-harness
A framework for few-shot evaluation of language models. observed · 2026-08-28
Health v2 · maintenance only
89/100
- Activity 99
- Release rhythm 71
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 54.0
- age_days: 2197
- days_rel: 114
- days_push: 7
- n_releases_24m: 11
Adoption not part of the score
13803 stars · 3518 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
A Python framework for few-shot evaluation of language models across hundreds of standard benchmarks (HellaSwag, MMLU, BIG-Bench, etc.). It supports HuggingFace, vLLM, SGLang, and API-hosted models with configurable tasks via YAML and a refactored CLI.
Use cases
- evaluate an LLM on standard benchmarks like MMLU or HellaSwag
- compare language models on few-shot tasks
- run the Open LLM Leaderboard task suite
- benchmark a HuggingFace or vLLM model
- evaluate models served via OpenAI-compatible APIs
- create custom evaluation tasks with YAML configs
- measure chain-of-thought reasoning performance
When to choose
- you need reproducible, standardized LLM benchmark scores
- you want to evaluate models across many backends (HF, vLLM, SGLang, APIs)
- you need configurable few-shot prompting and custom tasks
- you're submitting or reproducing Open LLM Leaderboard results
When to avoid
- you need broad multimodal (text+image) evaluation coverage - consider lmms-eval
- you want training or fine-tuning rather than evaluation
- you need a hosted no-code evaluation service
Facets
framework · maturity active
machine-learning llm-inference benchmarking testing cli large-language-models machine-learning developer-tools python cli windows llm-evaluation benchmarks few-shot huggingface vllm open-llm-leaderboard model-evaluation natural-language-processing gpu linux macos
4 sources
- readme: https://github.com/EleutherAI/lm-evaluation-harness · fetched 2026-08-28 · d2260704980a
- homepage: https://www.eleuther.ai · fetched 2026-08-29 · 46c7f1a08777
- site_page: https://www.eleuther.ai/about · fetched 2026-08-29 · 8a099f02f9a1
- site_page: https://www.eleuther.ai/releases · fetched 2026-08-29 · 6b771c6c6df0
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| EleutherAI/lm-evaluation-harness | main | 89 |
For agents
markdown · JSON · MCP: product_card(name="EleutherAI/lm-evaluation-harness")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem