# poloclub/diffusiondb

A large-scale text-to-image prompt gallery dataset based on Stable Diffusion

Repository: https://github.com/poloclub/diffusiondb
Canonical: https://ross.abutalabs.com/products/diffusiondb
Homepage: https://poloclub.github.io/diffusiondb
Language: Python
License: MIT
License Family: permissive
Topics: computer-vision, ai-art, image-generation, prompt-engineering, stable-diffusion
Last push: 2024-07-11T16:10:02+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 1423, "days_push": 783, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1390, forks 79 (observed 2026-08-28T04:04:35.586449+00:00)

## What it is
DiffusionDB is the first large-scale text-to-image prompt gallery dataset, containing 14 million images generated by Stable Diffusion from prompts and hyperparameters specified by real users. It is distributed via Hugging Face Datasets in two subsets (2M and Large) with parquet metadata tables and a Python downloader.

## Use cases
- study how prompts affect text-to-image generation quality
- train models to detect Stable Diffusion deepfakes
- build prompt engineering recommendation tools
- research human-AI interaction for generative models
- analyze real-world prompt and hyperparameter distributions
- train prompt expansion or prompt improvement models

## When to choose
- you need large-scale real-user prompts paired with generated images and generation parameters
- you are researching text-to-image model behavior, deepfake detection, or prompt optimization
- you want a citable, openly licensed dataset for generative AI research

## When to avoid
- you need images from models other than Stable Diffusion
- you cannot store multi-terabyte image collections (use the text-only metadata instead)
- you need curated or filtered commercial-grade training data with consent guarantees

## Facets
- artifact type: dataset
- maturity: stable
- function: machine-learning, stable-diffusion, prompt-engineering, data-science
- domain: artificial-intelligence, computer-vision, image-processing, large-language-models
- platform: python, cross-platform
- tags: text-to-image, diffusion-models, prompt-gallery, hugging-face, generative-ai, research-dataset

## Member repositories
- poloclub/diffusiondb (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:35.586449+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:39:39.875566+00:00, confidence not recorded.
  - readme: https://github.com/poloclub/diffusiondb (fetched 2026-08-28T04:04:35.586449+00:00, sha 6fc0e7860f36)
  - homepage: https://poloclub.github.io/diffusiondb (fetched 2026-08-29T11:54:48.827272+00:00, sha cfbe9a7567ba)
- Data as of 2026-08-30T08:39:29.467469+00:00.
