# vectara/hallucination-leaderboard

Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents

Repository: https://github.com/vectara/hallucination-leaderboard
Canonical: https://ross.abutalabs.com/products/hallucination-leaderboard
Homepage: https://vectara.com
Language: Python
License: Apache-2.0
License Family: permissive
Topics: generative-ai, hallucinations, llm
Last push: 2026-05-11T18:42:42+00:00

## Health v2 (maintenance only)
Score: 64/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 81, release rhythm 35, longevity 74
- inputs: {"age_days": 1037, "days_push": 114, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3306, forks 107 (observed 2026-08-28T04:07:55.854245+00:00)

## What it is
A public leaderboard maintained by Vectara that ranks LLMs by how often they hallucinate when summarizing short documents, computed with Vectara's HHEM hallucination evaluation model. It is updated regularly and also available as an interactive Hugging Face Space.

## Use cases
- compare llms for hallucination rates
- find which llm is most factually consistent at summarization
- check hallucination rate of a model before choosing it
- benchmark llms on factual consistency
- evaluate which model to use for document summarization
- track how llm hallucination rates change over time

## When to choose
- you need an up-to-date, third-party comparison of LLM factual consistency in summarization
- you are selecting a model for RAG or summarization pipelines where hallucination risk matters
- you want a simple ranked reference backed by a published evaluation model (HHEM)

## When to avoid
- you need task-level evaluation beyond short-document summarization
- you require a runnable evaluation harness for your own datasets rather than a published ranking
- you need benchmarks of reasoning, coding, or other non-summarization capabilities

## Facets
- artifact type: dataset
- maturity: active
- function: benchmarking, machine-learning, nlp
- domain: large-language-models, machine-learning, artificial-intelligence
- platform: python
- tags: llm-evaluation, hallucination-detection, leaderboard, summarization, factual-consistency, hhem, web

## Member repositories
- vectara/hallucination-leaderboard (main) score 64

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:55.854245+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:42:38.738695+00:00, confidence not recorded.
  - readme: https://github.com/vectara/hallucination-leaderboard (fetched 2026-08-28T04:07:55.854245+00:00, sha e5ae0435c5bd)
  - homepage: https://vectara.com (fetched 2026-08-29T09:35:28.768549+00:00, sha 0cda52f09586)
  - site_page: https://www.vectara.com/company/about-us (fetched 2026-08-29T09:35:28.779413+00:00, sha 26a119057c82)
  - site_page: https://docs.vectara.com (fetched 2026-08-29T09:35:28.781347+00:00, sha 2b6e4f00eb39)
  - site_page: https://docs.vectara.com/docs/api-reference/rest (fetched 2026-08-29T09:35:28.783129+00:00, sha 814dac386bed)
  - site_page: https://docs.vectara.com/docs/quickstart (fetched 2026-08-29T09:35:28.785131+00:00, sha 76bba62beee9)
  - site_page: https://www.vectara.com/developers/build/integrations (fetched 2026-08-29T09:35:28.787113+00:00, sha b315f01895b8)
  - site_page: https://www.vectara.com/faqs (fetched 2026-08-29T09:35:28.788793+00:00, sha b774cb51d360)
  - site_page: https://www.vectara.com/pricing (fetched 2026-08-29T09:35:28.777613+00:00, sha ccee5703a97c)
- Data as of 2026-08-30T08:39:29.467469+00:00.
