# centerforaisafety/hle

Humanity's Last Exam

Repository: https://github.com/centerforaisafety/hle
Canonical: https://ross.abutalabs.com/products/hle
Homepage: https://lastexam.ai
Language: Python
License: MIT
License Family: permissive
Last push: 2026-08-01T08:53:41+00:00

## Health v2 (maintenance only)
Score: 63/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 95, release rhythm 35, longevity 42
- inputs: {"age_days": 588, "days_push": 32, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1665, forks 109 (observed 2026-08-28T04:05:18.956263+00:00)

## What it is
Humanity's Last Exam (HLE) is a multi-modal benchmark of 2,500 expert-written questions across dozens of academic subjects, designed to test frontier AI models at the edge of human knowledge. The repository provides the dataset (hosted on Hugging Face) plus simple evaluation scripts for running models and grading their answers with an LLM judge.

## Use cases
- evaluate frontier LLMs on expert-level academic questions
- benchmark reasoning models across math, humanities, and sciences
- compare model accuracy and calibration on a hard closed-ended benchmark
- filter benchmark data out of training corpora using the canary string
- run automated grading of model answers with an LLM judge
- contribute new questions to the rolling HLE-Rolling version

## When to choose
- you need a challenging, broad-coverage benchmark for evaluating large language models
- you want automated, reproducible evaluation with multiple-choice and short-answer grading
- you need a citable, peer-reviewed benchmark (published in Nature) for research

## When to avoid
- you need a benchmark for everyday or beginner-level tasks
- you want a training dataset - the benchmark data must not appear in training corpora
- you need cheap, fast evaluation - frontier questions require large token budgets and LLM-judge calls

## Facets
- artifact type: dataset
- maturity: active
- function: benchmarking, llm-inference, machine-learning, data-science
- domain: artificial-intelligence, large-language-models, education, mathematics
- platform: python, cross-platform
- tags: llm-benchmark, evaluation, multimodal, academic-benchmark, huggingface-dataset, model-evaluation, natural-language-processing

## Member repositories
- centerforaisafety/hle (main) score 63

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:18.956263+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:43:01.483824+00:00, confidence not recorded.
  - readme: https://github.com/centerforaisafety/hle (fetched 2026-08-28T04:05:18.956263+00:00, sha fd3bd0049e1a)
  - homepage: https://lastexam.ai (fetched 2026-08-29T11:16:29.427275+00:00, sha 8492effff88c)
- Data as of 2026-08-30T08:39:29.467469+00:00.
