# Zjh-819/LLMDataHub

A quick guide (especially) for trending instruction finetuning datasets

Repository: https://github.com/Zjh-819/LLMDataHub
Canonical: https://ross.abutalabs.com/products/llmdatahub
License: MIT
License Family: permissive
Topics: chatbot, dataset, llm, chatgpt
Last push: 2023-11-28T09:41:28+00:00

## Health v2 (maintenance only)
Score: 30/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 88
- inputs: {"age_days": 1241, "days_push": 1009, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3411, forks 234 (observed 2026-08-28T04:08:03.436517+00:00)

## What it is
A curated catalog (awesome list) of open-source datasets for training large language models, covering alignment/SFT, domain-specific, pretraining, and multimodal corpora. Each entry includes links, size, language, usage type tags, and brief descriptions.

## Use cases
- find instruction finetuning datasets for llm training
- curated list of sft datasets for chatbot training
- where to get rlhf preference data
- pretraining corpora for open-source llms
- multimodal training datasets for vision-language models
- compare chatbot finetuning datasets by size and language

## When to choose
- you need to discover and compare open LLM training datasets before finetuning
- you want a regularly updated index of alignment, pretraining, and multimodal corpora
- you are researching dataset options for chatbot or instruction-following models

## When to avoid
- you need the actual dataset files or download tooling rather than links
- you want code or scripts for training, not a reference list
- you need guaranteed dataset quality evaluation rather than brief descriptions

## Facets
- artifact type: dataset
- maturity: active
- function: machine-learning, llm-training, data-science
- domain: large-language-models, machine-learning, chatbots, awesome-lists
- platform: cross-platform
- tags: awesome-list, instruction-tuning, sft, rlhf, pretraining, multimodal, curated-list, finetuning-datasets, natural-language-processing

## Member repositories
- Zjh-819/LLMDataHub (main) score 30

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:03.436517+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:38:30.687716+00:00, confidence not recorded.
  - readme: https://github.com/Zjh-819/LLMDataHub (fetched 2026-08-28T04:08:03.436517+00:00, sha 8a81b41dca9f)
- Data as of 2026-08-30T08:39:29.467469+00:00.
