Ross ROSS = Recommend OSS · open-source software intelligence for agents

centerforaisafety/hle resource

Humanity's Last Exam observed · 2026-08-28

github.com/centerforaisafety/hle · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

63/100

  • Activity 95
  • Release rhythm 35
  • Longevity 42

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 588
  • days_rel: n/a
  • days_push: 32
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1665 stars · 109 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Humanity's Last Exam (HLE) is a multi-modal benchmark of 2,500 expert-written questions across dozens of academic subjects, designed to test frontier AI models at the edge of human knowledge. The repository provides the dataset (hosted on Hugging Face) plus simple evaluation scripts for running models and grading their answers with an LLM judge.

Use cases

  • evaluate frontier LLMs on expert-level academic questions
  • benchmark reasoning models across math, humanities, and sciences
  • compare model accuracy and calibration on a hard closed-ended benchmark
  • filter benchmark data out of training corpora using the canary string
  • run automated grading of model answers with an LLM judge
  • contribute new questions to the rolling HLE-Rolling version

When to choose

  • you need a challenging, broad-coverage benchmark for evaluating large language models
  • you want automated, reproducible evaluation with multiple-choice and short-answer grading
  • you need a citable, peer-reviewed benchmark (published in Nature) for research

When to avoid

  • you need a benchmark for everyday or beginner-level tasks
  • you want a training dataset - the benchmark data must not appear in training corpora
  • you need cheap, fast evaluation - frontier questions require large token budgets and LLM-judge calls

Facets

dataset · maturity active

benchmarking llm-inference machine-learning data-science artificial-intelligence large-language-models education mathematics python cross-platform llm-benchmark evaluation multimodal academic-benchmark huggingface-dataset model-evaluation natural-language-processing

2 sources

Member repositories

RepositoryRoleHealth v2
centerforaisafety/hlemain63

For agents

markdown · JSON · MCP: product_card(name="centerforaisafety/hle")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem