# argilla-io/distilabel

Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

Repository: https://github.com/argilla-io/distilabel
Canonical: https://ross.abutalabs.com/products/distilabel
Homepage: https://distilabel.argilla.io
Language: Python
License: Apache-2.0
License Family: permissive
Topics: ai, huggingface, llms, openai, python, rlaif, rlhf, synthetic-data, synthetic-dataset-generation
Last push: 2026-08-24T22:15:36+00:00

## Health v2 (maintenance only)
Score: 74/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 40, longevity 75
- inputs: {"age_days": 1052, "days_push": 9, "days_rel": 582, "gap_med": 6.0, "n_releases_24m": 7}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3378, forks 255 (observed 2026-08-28T04:07:58.550683+00:00)

## What it is
Distilabel is a Python framework for building scalable pipelines that generate synthetic data and AI feedback, based on verified research papers. It supports both traditional NLP tasks and LLM scenarios like instruction following, dialogue generation, and LLM-as-a-judge workflows.

## Use cases
- generate synthetic instruction datasets for llm fine-tuning
- build rlaif pipelines with ai feedback
- create preference datasets for rlhf
- judge and label data with llm-as-a-judge
- generate synthetic training data from research papers
- scale up dataset generation pipelines with huggingface models

## When to choose
- you need reproducible, research-backed synthetic data pipelines
- you want to generate or judge datasets using multiple llm backends
- you are preparing instruction or preference datasets for fine-tuning

## When to avoid
- you need a no-code data labeling UI rather than programmatic pipelines
- your project requires actively maintained upstream support, since original authors have moved on
- you only need simple one-off prompt generation without pipeline orchestration

## Facets
- artifact type: framework
- maturity: maintenance
- function: data-generation, llm-training, rag, machine-learning, workflow-automation
- domain: artificial-intelligence, large-language-models, machine-learning
- platform: python, cross-platform
- tags: synthetic-data, ai-feedback, rlaif, rlhf, huggingface, data-pipelines, llm-judging, instruction-datasets, data-engineering, natural-language-processing

## Member repositories
- argilla-io/distilabel (main) score 74

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:58.550683+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:39:42.149923+00:00, confidence not recorded.
  - readme: https://github.com/argilla-io/distilabel (fetched 2026-08-28T04:07:58.550683+00:00, sha 52410af6e99d)
  - homepage: https://distilabel.argilla.io (fetched 2026-08-29T09:33:31.703209+00:00, sha 36c6c3c2e4f9)
- Data as of 2026-08-30T08:39:29.467469+00:00.
