# SophonPlus/ChineseNlpCorpus

搜集、整理、发布 中文 自然语言处理 语料/数据集，与 有志之士 共同 促进 中文 自然语言处理 的 发展。

Repository: https://github.com/SophonPlus/ChineseNlpCorpus
Canonical: https://ross.abutalabs.com/products/chinesenlpcorpus
Language: Jupyter Notebook
License Family: other
Last push: 2019-01-29T11:21:37+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 3084, "days_push": 2773, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 6594, forks 1422 (observed 2026-08-28T04:09:45.562324+00:00)

## What it is
A curated collection of Chinese natural language processing corpora and datasets, including sentiment analysis reviews, named entity recognition annotations, recommendation system ratings, and FAQ question-answering data. It serves as a central hub for downloading and exploring Chinese NLP datasets with introductory notebooks for each.

## Use cases
- find chinese sentiment analysis datasets
- download chinese named entity recognition training data
- get chinese question answering corpus for chatbot training
- find chinese review data for text classification
- obtain chinese recommendation system datasets
- train chinese nlp models on public corpora

## When to choose
- you need labeled Chinese text data for sentiment analysis or NER
- you are building or benchmarking Chinese NLP models
- you want FAQ-style question answering data in Chinese domains like finance or law

## When to avoid
- you need datasets in languages other than Chinese
- you need actively maintained datasets with guaranteed updates or a formal license
- you need English NLP corpora

## Facets
- artifact type: dataset
- maturity: maintenance
- function: nlp, machine-learning, data-science
- domain: machine-learning, data-science
- platform: cross-platform
- tags: chinese-nlp, corpus, sentiment-analysis, named-entity-recognition, question-answering, recommendation-systems, text-classification, natural-language-processing

## Member repositories
- SophonPlus/ChineseNlpCorpus (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:45.562324+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:43:33.160688+00:00, confidence not recorded.
  - readme: https://github.com/SophonPlus/ChineseNlpCorpus (fetched 2026-08-28T04:09:45.562324+00:00, sha b1bae6d5f600)
- Data as of 2026-08-30T08:39:29.467469+00:00.
