# MLGroupJLU/LLM-eval-survey

The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

Repository: https://github.com/MLGroupJLU/LLM-eval-survey
Canonical: https://ross.abutalabs.com/products/llm-eval-survey
Homepage: https://arxiv.org/abs/2307.03109
License Family: other
Topics: benchmark, evaluation, large-language-models, llm, llms, model-assessment
Last push: 2026-08-01T05:43:40+00:00

## Health v2 (maintenance only)
Score: 71/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 95, release rhythm 35, longevity 82
- inputs: {"age_days": 1158, "days_push": 32, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1609, forks 101 (observed 2026-08-28T04:05:10.858538+00:00)

## What it is
The official repository accompanying the survey paper 'A Survey on Evaluation of Large Language Models', curating papers and resources on LLM evaluation. It organizes research by what to evaluate (NLP, robustness, ethics, science, medicine, agents) and where/how to evaluate (benchmarks and methods).

## Use cases
- find papers on evaluating large language models
- survey of LLM benchmarks and evaluation methods
- learn how to assess LLM trustworthiness and robustness
- find benchmarks for medical or agent applications of LLMs
- keep up to date with new LLM evaluation research
- background reading for designing an LLM evaluation strategy

## When to choose
- you need a curated, organized reading list of LLM evaluation papers
- you are writing a related-work section on LLM assessment
- you want to discover benchmarks across many task domains

## When to avoid
- you need runnable evaluation code or a tool rather than a paper collection
- you need a single benchmark dataset to execute
- you require licensed software with support guarantees

## Facets
- artifact type: learning-resource
- maturity: active
- function: documentation, benchmarking
- domain: large-language-models, artificial-intelligence, tutorials, awesome-lists
- platform: -
- tags: survey-paper, llm-evaluation, benchmarks, paper-collection, reading-list, web-server

## Member repositories
- MLGroupJLU/LLM-eval-survey (main) score 71

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:10.858538+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:51:15.524786+00:00, confidence not recorded.
  - readme: https://github.com/MLGroupJLU/LLM-eval-survey (fetched 2026-08-28T04:05:10.858538+00:00, sha 4e9679b06e1a)
  - homepage: https://arxiv.org/abs/2307.03109 (fetched 2026-08-29T11:23:46.513894+00:00, sha 73a51ff63998)
  - site_page: https://info.arxiv.org/about/donate.html (fetched 2026-08-29T11:23:46.516804+00:00, sha cca9c3a11c56)
  - site_page: https://info.arxiv.org/about/ourmembers.html (fetched 2026-08-29T11:23:46.520797+00:00, sha 47cbc55ff1de)
  - site_page: https://info.arxiv.org/about (fetched 2026-08-29T11:23:46.523012+00:00, sha a1f16f915a9a)
  - site_page: https://info.arxiv.org/labs/index.html (fetched 2026-08-29T11:23:46.518923+00:00, sha b14a8d05a0ec)
- Data as of 2026-08-30T08:39:29.467469+00:00.
