datadreamer-dev/DataDreamer
DataDreamer: Prompt. Generate Synthetic Data. Train & Align Models. 🤖💤 observed · 2026-08-28
Health v2 · maintenance only
33/100
- Activity 4
- Release rhythm 40
- Longevity 84
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 0
- age_days: 1188
- days_rel: 577
- days_push: 577
- n_releases_24m: 8
Adoption not part of the score
1117 stars · 58 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
DataDreamer is an open-source Python library for prompting LLMs, generating synthetic datasets, and training or aligning models in reproducible workflows. It provides research-grade building blocks for data generation, distillation, instruction-tuning, and preference alignment.
Use cases
- generate synthetic training data with an llm
- distill gpt-4 capabilities into a smaller cheaper model
- instruction-tune a large language model
- align a model with human preferences
- augment an existing dataset using llms
- clean or filter a dataset with llm prompts
- bootstrap few-shot examples for prompting
- train a self-improving llm with self-rewarding
When to choose
- you need reproducible, research-grade synthetic data generation pipelines
- you want to distill a large LLM into a smaller model
- you are doing NLP/ML research involving instruction-tuning or alignment
- you want caching and saved outputs for LLM workflows out of the box
When to avoid
- you only need a simple OpenAI API wrapper without training features
- you need a production serving/inference platform rather than data and training workflows
- you work outside Python or the PyTorch/Transformers ecosystem
Facets
library · maturity active
machine-learning llm-training prompt-engineering rag data-generation etl machine-learning large-language-models data-science artificial-intelligence python cross-platform synthetic-data llm-workflows instruction-tuning alignment distillation pytorch transformers openai research-tools llmops natural-language-processing gpu
10 sources
- readme: https://github.com/datadreamer-dev/DataDreamer · fetched 2026-08-28 · 248c7326918f
- homepage: https://datadreamer.dev · fetched 2026-08-29 · f9eb22f1969b
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/installation.html · fetched 2026-08-29 · 44136fa355b3
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/quick_tour/index.html · fetched 2026-08-29 · a5da9539f76f
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/motivation_and_design.html · fetched 2026-08-29 · 35a2f166823a
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/quick_tour/abstract_to_tweet.html · fetched 2026-08-29 · aa515dd9acfd
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/quick_tour/attributed_prompts.html · fetched 2026-08-29 · 528c62dec0b1
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/quick_tour/openai_distillation.html · fetched 2026-08-29 · f455d2b07ca0
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/quick_tour/dataset_augmentation.html · fetched 2026-08-29 · f95b33c9b937
- site_page: https://datadreamer.dev/docs/latest/pages/get_started/quick_tour/dataset_cleaning.html · fetched 2026-08-29 · fc5a838745a8
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| datadreamer-dev/DataDreamer | main | 33 |
For agents
markdown · JSON · MCP: product_card(name="datadreamer-dev/DataDreamer")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem