# candlewill/Dialog_Corpus

用于训练中英文对话系统的语料库 Datasets for Training Chatbot System

Repository: https://github.com/candlewill/Dialog_Corpus
Canonical: https://ross.abutalabs.com/products/dialog_corpus
Language: Python
License Family: other
Topics: dataset, dialog, system, corpus, chatbot
Last push: 2020-09-23T21:06:45+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 3459, "days_push": 2170, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2056, forks 489 (observed 2026-08-28T04:06:09.893366+00:00)

## What it is
A curated collection of links to Chinese and English dialogue corpora for training chatbot systems, including movie subtitles, SMS corpora, QA pairs, and community chat logs. It is primarily an index of publicly available datasets rather than a software tool.

## Use cases
- find training data for a Chinese chatbot
- download movie subtitle dialogue corpora
- get QA pairs for a retrieval-based dialog system
- collect SMS conversation data for NLP research
- build a seq2seq conversation model with open datasets

## When to choose
- you need open dialog datasets for Chinese or English chatbot training
- you want a starting index of conversation corpora instead of scraping data yourself

## When to avoid
- you need a maintained, license-clear dataset with guaranteed quality
- you need actual data hosted in this repo rather than external links
- you need modern instruction-tuning or LLM training data

## Facets
- artifact type: dataset
- maturity: maintenance
- function: nlp, chatbot, machine-learning
- domain: chatbots, machine-learning
- platform: python, cross-platform
- tags: dialog-corpus, chinese-nlp, conversation-datasets, awesome-list, training-data, natural-language-processing

## Member repositories
- candlewill/Dialog_Corpus (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:09.893366+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:57:11.990067+00:00, confidence not recorded.
  - readme: https://github.com/candlewill/Dialog_Corpus (fetched 2026-08-28T04:06:09.893366+00:00, sha 03ce341fcaf6)
- Data as of 2026-08-30T08:39:29.467469+00:00.
