Ross ROSS = Recommend OSS · open-source software intelligence for agents

Zjh-819/LLMDataHub resource

A quick guide (especially) for trending instruction finetuning datasets observed · 2026-08-28

github.com/Zjh-819/LLMDataHub · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

30/100

  • Activity 0
  • Release rhythm 35
  • Longevity 88

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1241
  • days_rel: n/a
  • days_push: 1009
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

3411 stars · 234 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A curated catalog (awesome list) of open-source datasets for training large language models, covering alignment/SFT, domain-specific, pretraining, and multimodal corpora. Each entry includes links, size, language, usage type tags, and brief descriptions.

Use cases

  • find instruction finetuning datasets for llm training
  • curated list of sft datasets for chatbot training
  • where to get rlhf preference data
  • pretraining corpora for open-source llms
  • multimodal training datasets for vision-language models
  • compare chatbot finetuning datasets by size and language

When to choose

  • you need to discover and compare open LLM training datasets before finetuning
  • you want a regularly updated index of alignment, pretraining, and multimodal corpora
  • you are researching dataset options for chatbot or instruction-following models

When to avoid

  • you need the actual dataset files or download tooling rather than links
  • you want code or scripts for training, not a reference list
  • you need guaranteed dataset quality evaluation rather than brief descriptions

Facets

dataset · maturity active

machine-learning llm-training data-science large-language-models machine-learning chatbots awesome-lists cross-platform awesome-list instruction-tuning sft rlhf pretraining multimodal curated-list finetuning-datasets natural-language-processing

1 source

Member repositories

RepositoryRoleHealth v2
Zjh-819/LLMDataHubmain30

For agents

markdown · JSON · MCP: product_card(name="Zjh-819/LLMDataHub")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem