# hkust-nlp/ceval

Official github repo for C-Eval, a Chinese evaluation suite for foundation models [NeurIPS 2023]

Repository: https://github.com/hkust-nlp/ceval
Canonical: https://ross.abutalabs.com/products/ceval
Homepage: https://cevalbenchmark.com/
Language: Python
License: MIT
License Family: permissive
Last push: 2025-07-27T03:55:53+00:00

## Health v2 (maintenance only)
Score: 44/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 33, release rhythm 35, longevity 86
- inputs: {"age_days": 1209, "days_push": 402, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1867, forks 84 (observed 2026-08-28T04:05:46.342382+00:00)

## What it is
C-Eval is a comprehensive Chinese evaluation benchmark for foundation models, consisting of 13,948 multiple-choice questions across 52 disciplines and four difficulty levels. The repository provides the dataset, evaluation scripts, and leaderboards for measuring LLM performance on Chinese-language tasks.

## Use cases
- evaluate a large language model on Chinese-language knowledge
- benchmark my LLM against GPT-4 and ChatGPT on Chinese tasks
- find weaknesses of my model across 52 academic disciplines
- run standardized Chinese multiple-choice evaluations
- compare Chinese capabilities of open-source and proprietary models
- add a Chinese benchmark to my model evaluation pipeline

## When to choose
- you need a standardized, widely-cited Chinese benchmark for foundation models
- you want multi-discipline coverage from middle school to professional exams
- you want results comparable to published leaderboards and lm-evaluation-harness

## When to avoid
- you need evaluation in languages other than Chinese
- you need open-ended generation or reasoning benchmarks rather than multiple-choice
- you need a lightweight task-specific test rather than a broad academic suite

## Facets
- artifact type: dataset
- maturity: stable
- function: benchmarking, testing, data-generation
- domain: large-language-models, machine-learning, education
- platform: python
- tags: chinese-benchmark, evaluation-suite, multiple-choice, foundation-models, neurips-2023, llm-evaluation, natural-language-processing

## Member repositories
- hkust-nlp/ceval (main) score 44

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:46.342382+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:15:23.446160+00:00, confidence not recorded.
  - readme: https://github.com/hkust-nlp/ceval (fetched 2026-08-28T04:05:46.342382+00:00, sha 434c9d665c01)
- Data as of 2026-08-30T08:39:29.467469+00:00.
