Ross ROSS = Recommend OSS · open-source software intelligence for agents

liucongg/NLPDataSet resource

记录本人整理的一些数据集 observed · 2026-08-28

github.com/liucongg/NLPDataSet · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1856
  • days_rel: n/a
  • days_push: 1539
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1093 stars · 134 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A curated collection of Chinese NLP datasets gathered and cleaned by the author, covering named entity recognition (NER), text summarization, extractive reading comprehension (QA), and text similarity tasks. It consolidates 22 Chinese NER sources—spanning medical records, finance, e-commerce, social media, and news—into a unified BIO-tagged format.

Use cases

  • download chinese NER training data
  • chinese medical named entity recognition dataset from electronic medical records
  • chinese text similarity dataset like LCQMC
  • chinese extractive reading comprehension QA corpus
  • chinese abstractive summarization dataset
  • unified BIO-format chinese ner corpus combining MSRA CLUENER CMeEE
  • resume or finance domain NER dataset for model training

When to choose

  • you need many Chinese NER sources pre-cleaned into one consistent BIO-tagged format
  • you want a single aggregation of Chinese NLP benchmarks across medical, finance, e-commerce, and social media domains
  • you are training or benchmarking Chinese NER, QA, summarization, or similarity models and need ready-made data

When to avoid

  • you need English or multilingual datasets - everything here is Chinese
  • you require the latest official versions or exact per-source licensing - data was aggregated from third parties and last updated in mid-2022
  • you need nested entity annotations preserved - nested entities were flattened during BIO conversion
  • you need programmatic or API access - distribution is via a Baidu Netdisk download link

Facets

dataset · maturity stable

nlp machine-learning data-science machine-learning data-science chinese-nlp named-entity-recognition ner-dataset bio-format text-summarization reading-comprehension text-similarity medical-nlp dataset-collection corpus natural-language-processing

1 source

Member repositories

RepositoryRoleHealth v2
liucongg/NLPDataSetmain32

For agents

markdown · JSON · MCP: product_card(name="liucongg/NLPDataSet")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem