# wasiahmad/Awesome-LLM-Synthetic-Data

A reading list on LLM based Synthetic Data Generation 🔥

Repository: https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data
Canonical: https://ross.abutalabs.com/products/awesome-llm-synthetic-data
License: MIT
License Family: permissive
Last push: 2025-06-05T19:16:26+00:00

## Health v2 (maintenance only)
Score: 34/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 25, release rhythm 35, longevity 53
- inputs: {"age_days": 755, "days_push": 454, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1551, forks 95 (observed 2026-08-28T04:05:02.281414+00:00)

## What it is
A curated awesome-list of papers, tools, and blogs about synthetic data generation with large language models. It organizes resources by surveys, techniques, application areas (math reasoning, code generation, alignment, etc.), datasets, and tools.

## Use cases
- find papers on LLM synthetic data generation
- learn how to generate training data with LLMs
- research instruction tuning data generation methods
- find tools for generating synthetic datasets
- survey synthetic data for code generation or math reasoning
- find datasets generated by LLMs for model training

## When to choose
- you need a curated starting point for LLM synthetic data research
- you want recent papers and tools on data generation with LLMs
- you are surveying application areas like text-to-SQL or alignment data synthesis

## When to avoid
- you need runnable software rather than a reading list
- you want non-LLM synthetic data methods (GANs, simulation)
- you need a maintained library with APIs

## Facets
- artifact type: learning-resource
- maturity: active
- function: data-generation, llm-training, rag, prompt-engineering
- domain: large-language-models, machine-learning, artificial-intelligence, data-science, awesome-lists
- platform: -
- tags: awesome-list, synthetic-data, reading-list, papers, curated-resources, web-server

## Member repositories
- wasiahmad/Awesome-LLM-Synthetic-Data (main) score 34

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:02.281414+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:30:17.953503+00:00, confidence not recorded.
  - readme: https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data (fetched 2026-08-28T04:05:02.281414+00:00, sha b782173bf128)
- Data as of 2026-08-30T08:39:29.467469+00:00.
