Ross ROSS = Recommend OSS · open-source software intelligence for agents

vectara/hallucination-leaderboard resource

Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents observed · 2026-08-28

github.com/vectara/hallucination-leaderboard · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

64/100

  • Activity 81
  • Release rhythm 35
  • Longevity 74

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1037
  • days_rel: n/a
  • days_push: 114
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

3306 stars · 107 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A public leaderboard maintained by Vectara that ranks LLMs by how often they hallucinate when summarizing short documents, computed with Vectara's HHEM hallucination evaluation model. It is updated regularly and also available as an interactive Hugging Face Space.

Use cases

  • compare llms for hallucination rates
  • find which llm is most factually consistent at summarization
  • check hallucination rate of a model before choosing it
  • benchmark llms on factual consistency
  • evaluate which model to use for document summarization
  • track how llm hallucination rates change over time

When to choose

  • you need an up-to-date, third-party comparison of LLM factual consistency in summarization
  • you are selecting a model for RAG or summarization pipelines where hallucination risk matters
  • you want a simple ranked reference backed by a published evaluation model (HHEM)

When to avoid

  • you need task-level evaluation beyond short-document summarization
  • you require a runnable evaluation harness for your own datasets rather than a published ranking
  • you need benchmarks of reasoning, coding, or other non-summarization capabilities

Facets

dataset · maturity active

benchmarking machine-learning nlp large-language-models machine-learning artificial-intelligence python llm-evaluation hallucination-detection leaderboard summarization factual-consistency hhem web

9 sources

Member repositories

RepositoryRoleHealth v2
vectara/hallucination-leaderboardmain64

For agents

markdown · JSON · MCP: product_card(name="vectara/hallucination-leaderboard")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem