Ross ROSS = Recommend OSS · open-source software intelligence for agents

MLGroupJLU/LLM-eval-survey resource

The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models". observed · 2026-08-28

github.com/MLGroupJLU/LLM-eval-survey · homepage observed · 2026-08-28

Health v2 · maintenance only

71/100

  • Activity 95
  • Release rhythm 35
  • Longevity 82

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1158
  • days_rel: n/a
  • days_push: 32
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1609 stars · 101 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

The official repository accompanying the survey paper 'A Survey on Evaluation of Large Language Models', curating papers and resources on LLM evaluation. It organizes research by what to evaluate (NLP, robustness, ethics, science, medicine, agents) and where/how to evaluate (benchmarks and methods).

Use cases

  • find papers on evaluating large language models
  • survey of LLM benchmarks and evaluation methods
  • learn how to assess LLM trustworthiness and robustness
  • find benchmarks for medical or agent applications of LLMs
  • keep up to date with new LLM evaluation research
  • background reading for designing an LLM evaluation strategy

When to choose

  • you need a curated, organized reading list of LLM evaluation papers
  • you are writing a related-work section on LLM assessment
  • you want to discover benchmarks across many task domains

When to avoid

  • you need runnable evaluation code or a tool rather than a paper collection
  • you need a single benchmark dataset to execute
  • you require licensed software with support guarantees

Facets

learning-resource · maturity active

documentation benchmarking large-language-models artificial-intelligence tutorials awesome-lists survey-paper llm-evaluation benchmarks paper-collection reading-list web-server

6 sources

Member repositories

RepositoryRoleHealth v2
MLGroupJLU/LLM-eval-surveymain71

For agents

markdown · JSON · MCP: product_card(name="MLGroupJLU/LLM-eval-survey")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem