# Toyhom/Chinese-medical-dialogue-data

Chinese medical dialogue data 中文医疗对话数据集

Repository: https://github.com/Toyhom/Chinese-medical-dialogue-data
Canonical: https://ross.abutalabs.com/products/chinese-medical-dialogue-data
Language: Python
License: MIT
License Family: permissive
Last push: 2023-08-18T05:34:28+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2459, "days_push": 1111, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1756, forks 295 (observed 2026-08-28T04:05:32.158147+00:00)

## What it is
A Chinese medical dialogue dataset containing 792,099 patient-doctor question-answer pairs organized into six departments: internal medicine, surgery, pediatrics, oncology, obstetrics/gynecology, and andrology. The data ships as CSV files (department, title, question, answer) and includes instruction-tuning format examples prepared for fine-tuning models such as ChatGLM-6B.

## Use cases
- find a Chinese medical dialogue dataset for NLP research
- fine-tune ChatGLM-6B on medical QA data
- train a Chinese medical chatbot
- get a doctor-patient question-answer corpus in Chinese
- instruction-tuning data for a medical LLM assistant
- benchmark LoRA or P-Tuning on domain-specific Chinese text

## When to choose
- You need large-scale Chinese medical QA pairs for LLM fine-tuning or evaluation
- You want ready-made instruction/input/output JSON for ChatGLM-style training
- You want a permissively licensed (MIT) dataset covering six medical departments
- You are reproducing or comparing LoRA / P-Tuning V2 results on medical dialogue

## When to avoid
- You need English or multilingual medical dialogue data
- You require clinically verified or expert-reviewed medical answers, since content originates from online consultation replies
- You need an application or API for medical QA rather than raw training data
- Your project needs actively maintained or updated data

## Facets
- artifact type: dataset
- maturity: stable
- function: machine-learning, llm-training, nlp
- domain: healthcare, large-language-models, machine-learning
- platform: python
- tags: chinese-nlp, medical-qa, dialogue-dataset, fine-tuning-data, chatglm, question-answering, csv-dataset, medical-chatbot, natural-language-processing

## Member repositories
- Toyhom/Chinese-medical-dialogue-data (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:32.158147+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:27:50.978614+00:00, confidence not recorded.
  - readme: https://github.com/Toyhom/Chinese-medical-dialogue-data (fetched 2026-08-28T04:05:32.158147+00:00, sha 5030c5671d29)
- Data as of 2026-08-30T08:39:29.467469+00:00.
