esbatmop/MNBVC resource
MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。 observed · 2026-08-28
Health v2 · maintenance only
75/100
- Activity 97
- Release rhythm 35
- Longevity 95
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1341
- days_rel: n/a
- days_push: 18
- n_releases_24m: 0
Adoption not part of the score
4267 stars · 296 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
MNBVC is a massive, continuously growing open-source Chinese text corpus aiming to rival the scale of data used to train ChatGPT, including news, novels, wiki, chat logs, poetry, and niche internet culture. The project also provides companion data-cleaning, OCR, deduplication, and code-repository crawling tools.
Use cases
- download a large-scale Chinese corpus for LLM pretraining
- find diverse Chinese text data including niche internet slang
- get cleaned Chinese datasets from Hugging Face or ModelScope
- clean and deduplicate large Chinese text corpora
- extract text from PDFs and images for training data
- crawl GitHub repositories to build code training corpora
When to choose
- you need tens of terabytes of Chinese text for training or fine-tuning language models
- you want broad coverage of Chinese internet culture including minority communities
- you need companion tools for Chinese encoding detection, deduplication, and corpus cleaning
When to avoid
- you need curated, indexed, or copyright-cleared data since the project performs no copyright review
- you need small, well-annotated datasets for supervised tasks rather than raw pretraining text
- you work with non-Chinese languages primarily
Facets
dataset · maturity active
nlp machine-learning llm-training ocr data-science etl large-language-models machine-learning python cross-platform chinese-corpus text-corpus llm-pretraining data-cleaning web-scraping multimodal open-dataset natural-language-processing data-engineering
1 source
- readme: https://github.com/esbatmop/MNBVC · fetched 2026-08-28 · 9be53906a57d
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| esbatmop/MNBVC | main | 75 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem