centerforaisafety/hle resource
Humanity's Last Exam observed · 2026-08-28
Health v2 · maintenance only
63/100
- Activity 95
- Release rhythm 35
- Longevity 42
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 588
- days_rel: n/a
- days_push: 32
- n_releases_24m: 0
Adoption not part of the score
1665 stars · 109 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Humanity's Last Exam (HLE) is a multi-modal benchmark of 2,500 expert-written questions across dozens of academic subjects, designed to test frontier AI models at the edge of human knowledge. The repository provides the dataset (hosted on Hugging Face) plus simple evaluation scripts for running models and grading their answers with an LLM judge.
Use cases
- evaluate frontier LLMs on expert-level academic questions
- benchmark reasoning models across math, humanities, and sciences
- compare model accuracy and calibration on a hard closed-ended benchmark
- filter benchmark data out of training corpora using the canary string
- run automated grading of model answers with an LLM judge
- contribute new questions to the rolling HLE-Rolling version
When to choose
- you need a challenging, broad-coverage benchmark for evaluating large language models
- you want automated, reproducible evaluation with multiple-choice and short-answer grading
- you need a citable, peer-reviewed benchmark (published in Nature) for research
When to avoid
- you need a benchmark for everyday or beginner-level tasks
- you want a training dataset - the benchmark data must not appear in training corpora
- you need cheap, fast evaluation - frontier questions require large token budgets and LLM-judge calls
Facets
dataset · maturity active
benchmarking llm-inference machine-learning data-science artificial-intelligence large-language-models education mathematics python cross-platform llm-benchmark evaluation multimodal academic-benchmark huggingface-dataset model-evaluation natural-language-processing
2 sources
- readme: https://github.com/centerforaisafety/hle · fetched 2026-08-28 · fd3bd0049e1a
- homepage: https://lastexam.ai · fetched 2026-08-29 · 8492effff88c
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| centerforaisafety/hle | main | 63 |
For agents
markdown · JSON · MCP: product_card(name="centerforaisafety/hle")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem