# esbatmop/MNBVC

MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化，也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。

Repository: https://github.com/esbatmop/MNBVC
Canonical: https://ross.abutalabs.com/products/mnbvc
License: MIT
License Family: permissive
Topics: chinese, chinese-language, chinese-nlp, chinese-simplified, corpus-data, nlp, nlp-machine-learning
Last push: 2026-08-15T11:52:11+00:00

## Health v2 (maintenance only)
Score: 75/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 97, release rhythm 35, longevity 95
- inputs: {"age_days": 1341, "days_push": 18, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 4267, forks 296 (observed 2026-08-28T04:08:40.732789+00:00)

## What it is
MNBVC is a massive, continuously growing open-source Chinese text corpus aiming to rival the scale of data used to train ChatGPT, including news, novels, wiki, chat logs, poetry, and niche internet culture. The project also provides companion data-cleaning, OCR, deduplication, and code-repository crawling tools.

## Use cases
- download a large-scale Chinese corpus for LLM pretraining
- find diverse Chinese text data including niche internet slang
- get cleaned Chinese datasets from Hugging Face or ModelScope
- clean and deduplicate large Chinese text corpora
- extract text from PDFs and images for training data
- crawl GitHub repositories to build code training corpora

## When to choose
- you need tens of terabytes of Chinese text for training or fine-tuning language models
- you want broad coverage of Chinese internet culture including minority communities
- you need companion tools for Chinese encoding detection, deduplication, and corpus cleaning

## When to avoid
- you need curated, indexed, or copyright-cleared data since the project performs no copyright review
- you need small, well-annotated datasets for supervised tasks rather than raw pretraining text
- you work with non-Chinese languages primarily

## Facets
- artifact type: dataset
- maturity: active
- function: nlp, machine-learning, llm-training, ocr, data-science, etl
- domain: large-language-models, machine-learning
- platform: python, cross-platform
- tags: chinese-corpus, text-corpus, llm-pretraining, data-cleaning, web-scraping, multimodal, open-dataset, natural-language-processing, data-engineering

## Member repositories
- esbatmop/MNBVC (main) score 75

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:40.732789+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:22:02.716688+00:00, confidence not recorded.
  - readme: https://github.com/esbatmop/MNBVC (fetched 2026-08-28T04:08:40.732789+00:00, sha 9be53906a57d)
- Data as of 2026-08-30T08:39:29.467469+00:00.
