Ross ROSS = Recommend OSS · open-source software intelligence for agents

CLUEbenchmark/CLUEDatasetSearch resource

搜索所有中文NLP数据集,附常用英文NLP数据集 observed · 2026-08-28

github.com/CLUEbenchmark/CLUEDatasetSearch · homepage · Python observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2385
  • days_rel: n/a
  • days_push: 1381
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

4454 stars · 627 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A curated searchable catalog of Chinese NLP datasets (with common English NLP datasets included), organized by task such as NER, QA, text classification, matching, summarization, machine translation, and knowledge graphs. It is part of the CLUE benchmark ecosystem and links to dataset sources rather than hosting data itself.

Use cases

  • find Chinese NER datasets
  • locate Chinese question answering datasets for training
  • search text classification datasets in Chinese
  • find Chinese machine translation corpora
  • discover sentiment analysis datasets for Chinese
  • find reading comprehension datasets in Chinese
  • compare Chinese and English NLP datasets

When to choose

  • you need Chinese-language NLP training or evaluation data and want a task-organized index
  • you are benchmarking Chinese language understanding models and need dataset references
  • you want a quick catalog of both Chinese and common English NLP datasets

When to avoid

  • you need the actual dataset files hosted and maintained in one place - it only links to external sources
  • you need datasets for languages other than Chinese and English
  • you need a programmatically queryable dataset API rather than a curated list

Facets

dataset · maturity maintenance

nlp search-engine data-science machine-learning data-science localization python cross-platform chinese-nlp dataset-catalog ner question-answering text-classification machine-translation knowledge-graph sentiment-analysis text-similarity text-summarization reading-comprehension corpus awesome-list natural-language-processing

3 sources

Member repositories

RepositoryRoleHealth v2
CLUEbenchmark/CLUEDatasetSearchmain32

For agents

markdown · JSON · MCP: product_card(name="CLUEbenchmark/CLUEDatasetSearch")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem