# CLUEbenchmark/CLUEDatasetSearch

搜索所有中文NLP数据集，附常用英文NLP数据集

Repository: https://github.com/CLUEbenchmark/CLUEDatasetSearch
Canonical: https://ross.abutalabs.com/products/cluedatasetsearch
Homepage: https://www.cluebenchmarks.com/dataSet_search.html
Language: Python
License Family: other
Topics: nlp, datasets, chinese, ner, qa, match, text-classification, machine-translation, knowledge-graph, corpus, machine-reading-comprehension, sentiment-analysis, text-similarity, text-summarization
Last push: 2022-11-21T08:04:52+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2385, "days_push": 1381, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 4454, forks 627 (observed 2026-08-28T04:08:50.543654+00:00)

## What it is
A curated searchable catalog of Chinese NLP datasets (with common English NLP datasets included), organized by task such as NER, QA, text classification, matching, summarization, machine translation, and knowledge graphs. It is part of the CLUE benchmark ecosystem and links to dataset sources rather than hosting data itself.

## Use cases
- find Chinese NER datasets
- locate Chinese question answering datasets for training
- search text classification datasets in Chinese
- find Chinese machine translation corpora
- discover sentiment analysis datasets for Chinese
- find reading comprehension datasets in Chinese
- compare Chinese and English NLP datasets

## When to choose
- you need Chinese-language NLP training or evaluation data and want a task-organized index
- you are benchmarking Chinese language understanding models and need dataset references
- you want a quick catalog of both Chinese and common English NLP datasets

## When to avoid
- you need the actual dataset files hosted and maintained in one place - it only links to external sources
- you need datasets for languages other than Chinese and English
- you need a programmatically queryable dataset API rather than a curated list

## Facets
- artifact type: dataset
- maturity: maintenance
- function: nlp, search-engine, data-science
- domain: machine-learning, data-science, localization
- platform: python, cross-platform
- tags: chinese-nlp, dataset-catalog, ner, question-answering, text-classification, machine-translation, knowledge-graph, sentiment-analysis, text-similarity, text-summarization, reading-comprehension, corpus, awesome-list, natural-language-processing

## Member repositories
- CLUEbenchmark/CLUEDatasetSearch (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:50.543654+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:20:44.990303+00:00, confidence not recorded.
  - readme: https://github.com/CLUEbenchmark/CLUEDatasetSearch (fetched 2026-08-28T04:08:50.543654+00:00, sha 293da1225f0d)
  - homepage: https://www.cluebenchmarks.com/dataSet_search.html (fetched 2026-08-29T09:07:37.956305+00:00, sha 0f3d32cd5926)
  - site_page: https://www.cluebenchmarks.com/aboutClue.html (fetched 2026-08-29T09:07:37.965679+00:00, sha b318c2db06a4)
- Data as of 2026-08-30T08:39:29.467469+00:00.
