tatsu-lab/alpaca_eval
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast. observed · 2026-08-28
Health v2 · maintenance only
36/100
- Activity 36
- Release rhythm 8
- Longevity 85
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1196
- days_rel: 614
- days_push: 389
- n_releases_24m: 1
Adoption not part of the score
2012 stars · 315 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
AlpacaEval is an LLM-based automatic evaluation framework for instruction-following language models, producing win rates against a GPT-4 baseline that correlate strongly (0.98) with human preferences from ChatBot Arena. It includes a public leaderboard, evaluation sets, GPT-4-based auto-annotators, and length-controlled win-rate metrics.
Use cases
- evaluate an instruction-following LLM automatically
- compare chat LLMs on a leaderboard cheaply and fast
- measure win rates of my model against GPT-4
- build a custom LLM leaderboard
- create a new automatic evaluator for LLM outputs
- benchmark chat models without human annotation
- check how well my model follows user instructions
When to choose
- you need a fast, cheap, human-correlated benchmark for chat/instruction-following LLMs
- you want to submit a model to a recognized community leaderboard
- you want to develop or validate new automatic evaluators or eval sets
When to avoid
- you need comprehensive safety or capability testing beyond general instruction following
- you cannot or will not use OpenAI API credits for annotation
- you require a gold-standard human evaluation rather than an LLM-based proxy
Facets
library · maturity active
benchmarking llm-inference machine-learning cli large-language-models machine-learning developer-tools python cli cross-platform llm-evaluation leaderboard instruction-following rlhf gpt-4-annotator win-rate chatbot-benchmark natural-language-processing
2 sources
- readme: https://github.com/tatsu-lab/alpaca_eval · fetched 2026-08-28 · 7f3fc1e48756
- homepage: https://tatsu-lab.github.io/alpaca_eval/ · fetched 2026-08-29 · ef7dda755582
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| tatsu-lab/alpaca_eval | main | 36 |
For agents
markdown · JSON · MCP: product_card(name="tatsu-lab/alpaca_eval")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem