Ross ROSS = Recommend OSS · open-source software intelligence for agents

jeinlee1991/chinese-llm-benchmark resource

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。 observed · 2026-08-28

github.com/jeinlee1991/chinese-llm-benchmark · homepage observed · 2026-08-28

Health v2 · maintenance only

89/100

  • Activity 99
  • Release rhythm 80
  • Longevity 84

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 5.0
  • age_days: 1186
  • days_rel: 134
  • days_push: 10
  • n_releases_24m: 57

Full methodology

Adoption not part of the score

6401 stars · 261 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

ReLE (formerly CLiB) is a continuously updated Chinese-language LLM capability benchmark and leaderboard covering ~400 commercial and open-source models across 7 domains and ~300 fine-grained dimensions (education, healthcare, finance, legal, reasoning, language, agent/tool use). It also publishes a defect (badcase) database of over 2 million cases for community research and model improvement.

Use cases

  • compare chinese llm leaderboard rankings
  • evaluate which chinese large language model to choose
  • find weaknesses and failure cases of llms
  • benchmark llm performance in education or medical exams
  • compare open-source vs commercial chinese models
  • research llm hallucination and refusal cases
  • evaluate agent and tool-calling ability of chinese models

When to choose

  • you need up-to-date, live evaluations of Chinese-language LLMs
  • you want domain-specific rankings (education, healthcare, finance, legal, reasoning)
  • you need a large badcase/defect corpus for model analysis or improvement
  • you are selecting between Chinese commercial and open-source models

When to avoid

  • you need English-only or multilingual benchmarking
  • you need a runnable evaluation harness for your own private test sets rather than published rankings
  • you require a formally licensed, reproducible academic benchmark (no license specified)

Facets

dataset · maturity active

benchmarking llm-inference machine-learning data-science large-language-models artificial-intelligence education healthcare fintech legal python cross-platform llm-evaluation chinese-llm leaderboard benchmark badcase-database model-selection rele live-evaluation natural-language-processing web

3 sources

Member repositories

RepositoryRoleHealth v2
jeinlee1991/chinese-llm-benchmarkmain89

For agents

markdown · JSON · MCP: product_card(name="jeinlee1991/chinese-llm-benchmark")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem