CLUEbenchmark/CLUEDatasetSearch resource
搜索所有中文NLP数据集,附常用英文NLP数据集 observed · 2026-08-28
Health v2 · maintenance only
32/100
- Activity 0
- Release rhythm 35
- Longevity 100
Flags: no_releases no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 2385
- days_rel: n/a
- days_push: 1381
- n_releases_24m: 0
Adoption not part of the score
4454 stars · 627 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
A curated searchable catalog of Chinese NLP datasets (with common English NLP datasets included), organized by task such as NER, QA, text classification, matching, summarization, machine translation, and knowledge graphs. It is part of the CLUE benchmark ecosystem and links to dataset sources rather than hosting data itself.
Use cases
- find Chinese NER datasets
- locate Chinese question answering datasets for training
- search text classification datasets in Chinese
- find Chinese machine translation corpora
- discover sentiment analysis datasets for Chinese
- find reading comprehension datasets in Chinese
- compare Chinese and English NLP datasets
When to choose
- you need Chinese-language NLP training or evaluation data and want a task-organized index
- you are benchmarking Chinese language understanding models and need dataset references
- you want a quick catalog of both Chinese and common English NLP datasets
When to avoid
- you need the actual dataset files hosted and maintained in one place - it only links to external sources
- you need datasets for languages other than Chinese and English
- you need a programmatically queryable dataset API rather than a curated list
Facets
dataset · maturity maintenance
nlp search-engine data-science machine-learning data-science localization python cross-platform chinese-nlp dataset-catalog ner question-answering text-classification machine-translation knowledge-graph sentiment-analysis text-similarity text-summarization reading-comprehension corpus awesome-list natural-language-processing
3 sources
- readme: https://github.com/CLUEbenchmark/CLUEDatasetSearch · fetched 2026-08-28 · 293da1225f0d
- homepage: https://www.cluebenchmarks.com/dataSet_search.html · fetched 2026-08-29 · 0f3d32cd5926
- site_page: https://www.cluebenchmarks.com/aboutClue.html · fetched 2026-08-29 · b318c2db06a4
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| CLUEbenchmark/CLUEDatasetSearch | main | 32 |
For agents
markdown · JSON · MCP: product_card(name="CLUEbenchmark/CLUEDatasetSearch")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem