# jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测：中文AI大模型能力评测（持续更新）：目前已囊括374个大模型，覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型， 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜，也提供规模超200万的大模型缺陷库！方便广大社区研究分析、改进大模型。

Repository: https://github.com/jeinlee1991/chinese-llm-benchmark
Canonical: https://ross.abutalabs.com/products/chinese-llm-benchmark
Homepage: https://nonelinear.com
License Family: other
Topics: agentic-ai, artificial-intelligence, llm-agent, llm-evaluation
Last push: 2026-08-23T10:20:24+00:00

## Health v2 (maintenance only)
Score: 89/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 99, release rhythm 80, longevity 84
- inputs: {"age_days": 1186, "days_push": 10, "days_rel": 134, "gap_med": 5.0, "n_releases_24m": 57}
- flags: no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 6401, forks 261 (observed 2026-08-28T04:09:43.100203+00:00)

## What it is
ReLE (formerly CLiB) is a continuously updated Chinese-language LLM capability benchmark and leaderboard covering ~400 commercial and open-source models across 7 domains and ~300 fine-grained dimensions (education, healthcare, finance, legal, reasoning, language, agent/tool use). It also publishes a defect (badcase) database of over 2 million cases for community research and model improvement.

## Use cases
- compare chinese llm leaderboard rankings
- evaluate which chinese large language model to choose
- find weaknesses and failure cases of llms
- benchmark llm performance in education or medical exams
- compare open-source vs commercial chinese models
- research llm hallucination and refusal cases
- evaluate agent and tool-calling ability of chinese models

## When to choose
- you need up-to-date, live evaluations of Chinese-language LLMs
- you want domain-specific rankings (education, healthcare, finance, legal, reasoning)
- you need a large badcase/defect corpus for model analysis or improvement
- you are selecting between Chinese commercial and open-source models

## When to avoid
- you need English-only or multilingual benchmarking
- you need a runnable evaluation harness for your own private test sets rather than published rankings
- you require a formally licensed, reproducible academic benchmark (no license specified)

## Facets
- artifact type: dataset
- maturity: active
- function: benchmarking, llm-inference, machine-learning, data-science
- domain: large-language-models, artificial-intelligence, education, healthcare, fintech, legal
- platform: python, cross-platform
- tags: llm-evaluation, chinese-llm, leaderboard, benchmark, badcase-database, model-selection, rele, live-evaluation, natural-language-processing, web

## Member repositories
- jeinlee1991/chinese-llm-benchmark (main) score 89

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:43.100203+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:45:00.954459+00:00, confidence not recorded.
  - readme: https://github.com/jeinlee1991/chinese-llm-benchmark (fetched 2026-08-28T04:09:43.100203+00:00, sha 71233a1b7aec)
  - homepage: https://nonelinear.com (fetched 2026-08-29T08:42:11.093204+00:00, sha d57863145cbe)
  - site_page: https://nonelinear.com/static/about.html (fetched 2026-08-29T08:42:11.102359+00:00, sha 70d21e29ca2d)
- Data as of 2026-08-30T08:39:29.467469+00:00.
