# thunlp/UltraChat

Large-scale, Informative, and Diverse Multi-round Chat Data (and Models)

Repository: https://github.com/thunlp/UltraChat
Canonical: https://ross.abutalabs.com/products/ultrachat
Language: Python
License: MIT
License Family: permissive
Topics: large-language-models, chatbot, deep-learning, chatgpt
Last push: 2024-03-13T02:49:52+00:00

## Health v2 (maintenance only)
Score: 30/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 89
- inputs: {"age_days": 1248, "days_push": 903, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2890, forks 144 (observed 2026-08-28T04:07:28.328883+00:00)

## What it is
UltraChat is a large-scale dataset of 1.57M informative, diverse multi-round dialogue instructions generated with LLMs, plus UltraLM, a series of chat models fine-tuned on it. It is primarily a research dataset for training open-source instruction-following chat models.

## Use cases
- download multi-round chat instruction data for fine-tuning an LLM
- train an instruction-following chatbot with SFT
- find open dialogue datasets for supervised fine-tuning
- reproduce open-source chat models like UltraLM
- get instruction tuning data as an alternative to ShareGPT

## When to choose
- you need large-scale multi-round instruction data for supervised fine-tuning
- you want a permissively licensed (MIT) chat dataset for LLM research
- you want a proven dataset behind top AlpacaEval open models

## When to avoid
- you need preference/RLHF data rather than SFT dialogues (use UltraFeedback instead)
- you want a production chatbot application rather than training data
- you need human-written rather than model-generated dialogues

## Facets
- artifact type: dataset
- maturity: maintenance
- function: machine-learning, llm-training, data-generation
- domain: large-language-models, chatbots, deep-learning, machine-learning
- platform: python
- tags: instruction-tuning, chat-data, sft, llama, dialogue-dataset

## Member repositories
- thunlp/UltraChat (main) score 30

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:28.328883+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:35:20.698693+00:00, confidence not recorded.
  - readme: https://github.com/thunlp/UltraChat (fetched 2026-08-28T04:07:28.328883+00:00, sha 6ac501028066)
- Data as of 2026-08-30T08:39:29.467469+00:00.
