liucongg/NLPDataSet resource
记录本人整理的一些数据集 observed · 2026-08-28
Health v2 · maintenance only
32/100
- Activity 0
- Release rhythm 35
- Longevity 100
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1856
- days_rel: n/a
- days_push: 1539
- n_releases_24m: 0
Adoption not part of the score
1093 stars · 134 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A curated collection of Chinese NLP datasets gathered and cleaned by the author, covering named entity recognition (NER), text summarization, extractive reading comprehension (QA), and text similarity tasks. It consolidates 22 Chinese NER sources—spanning medical records, finance, e-commerce, social media, and news—into a unified BIO-tagged format.
Use cases
- download chinese NER training data
- chinese medical named entity recognition dataset from electronic medical records
- chinese text similarity dataset like LCQMC
- chinese extractive reading comprehension QA corpus
- chinese abstractive summarization dataset
- unified BIO-format chinese ner corpus combining MSRA CLUENER CMeEE
- resume or finance domain NER dataset for model training
When to choose
- you need many Chinese NER sources pre-cleaned into one consistent BIO-tagged format
- you want a single aggregation of Chinese NLP benchmarks across medical, finance, e-commerce, and social media domains
- you are training or benchmarking Chinese NER, QA, summarization, or similarity models and need ready-made data
When to avoid
- you need English or multilingual datasets - everything here is Chinese
- you require the latest official versions or exact per-source licensing - data was aggregated from third parties and last updated in mid-2022
- you need nested entity annotations preserved - nested entities were flattened during BIO conversion
- you need programmatic or API access - distribution is via a Baidu Netdisk download link
Facets
dataset · maturity stable
nlp machine-learning data-science machine-learning data-science chinese-nlp named-entity-recognition ner-dataset bio-format text-summarization reading-comprehension text-similarity medical-nlp dataset-collection corpus natural-language-processing
1 source
- readme: https://github.com/liucongg/NLPDataSet · fetched 2026-08-28 · 8cbdd754ce44
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| liucongg/NLPDataSet | main | 32 |
For agents
markdown · JSON · MCP: product_card(name="liucongg/NLPDataSet")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem