# hikariming/chat-dataset-baseline

人工精调的中文对话数据集和一段chatglm的微调代码

Repository: https://github.com/hikariming/chat-dataset-baseline
Canonical: https://ross.abutalabs.com/products/chat-dataset-baseline
Language: Jupyter Notebook
License Family: other
Topics: alpaca, chatglm, dataset
Last push: 2025-05-03T02:36:56+00:00

## Health v2 (maintenance only)
Score: 39/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 19, release rhythm 35, longevity 90
- inputs: {"age_days": 1265, "days_push": 488, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1191, forks 97 (observed 2026-08-28T04:03:56.169545+00:00)

## What it is
A curated Chinese conversational dataset plus fine-tuning scripts for training chat models like ChatGLM, built on top of LLaMA-Factory. It includes preprocessing code to customize model identity and training scripts for LoRA or full-parameter SFT.

## Use cases
- fine-tune a Chinese chat model
- get a curated Chinese instruction dataset
- train ChatGLM with LoRA
- create a custom Chinese assistant baseline
- translate and adapt alpaca-style data for Chinese

## When to choose
- You want a ready-made Chinese dialogue dataset for SFT
- You use LLaMA-Factory and want drop-in data and training scripts
- You need a base Chinese chat model to further customize for a domain

## When to avoid
- You need English or multilingual datasets
- You want an actively maintained project (author recommends LLaMA-Factory instead)
- You need a licensed dataset (no license specified)

## Facets
- artifact type: dataset
- maturity: maintenance
- function: machine-learning, llm-training, data-generation
- domain: large-language-models, machine-learning
- platform: python
- tags: chinese-dataset, chatglm, alpaca, sft, lora, llama-factory, fine-tuning, natural-language-processing

## Member repositories
- hikariming/chat-dataset-baseline (main) score 39

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:56.169545+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:22:42.332573+00:00, confidence not recorded.
  - readme: https://github.com/hikariming/chat-dataset-baseline (fetched 2026-08-28T04:03:56.169545+00:00, sha c05478f7589b)
- Data as of 2026-08-30T08:39:29.467469+00:00.
