YangLing0818/RPG-DiffusionMaster
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG) observed · 2026-08-28
Health v2 · maintenance only
28/100
- Activity 4
- Release rhythm 35
- Longevity 68
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 955
- days_rel: n/a
- days_push: 578
- n_releases_24m: 0
Adoption not part of the score
1842 stars · 100 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Official implementation of RPG (Recaption, Plan, Generate), a training-free framework that uses multimodal LLMs as prompt recaptioners and region planners to improve compositional text-to-image generation and editing with diffusion models. It supports proprietary and open-source MLLMs and arbitrary diffusion backbones, enabling high-resolution, multi-object image generation.
Use cases
- generate complex multi-object images from long text prompts
- improve compositional text-to-image generation with diffusion models
- edit images with text-guided closed-loop generation
- generate super high resolution images with regional diffusion
- use GPT-4 or local MLLMs to plan image subregions
- research training-free prompt planning for diffusion models
When to choose
- you need better multi-object composition and text-image alignment than vanilla Stable Diffusion or SDXL
- you want a training-free framework compatible with various MLLMs and diffusion backbones
- you need high-resolution or region-wise controlled image generation and editing
When to avoid
- you need a production-ready end-user image generation app rather than research code
- you cannot access an MLLM API or run a local multimodal model
- you need fast single-prompt generation without LLM planning overhead
Facets
library · maturity active
llm-inference image-processing prompt-engineering machine-learning artificial-intelligence image-processing large-language-models deep-learning python cross-platform text-to-image diffusion-models multimodal-llm stable-diffusion image-editing research-code icml-2024 training-free gpu
3 sources
- readme: https://github.com/YangLing0818/RPG-DiffusionMaster · fetched 2026-08-28 · 16ff95d54381
- homepage: https://proceedings.mlr.press/v235/yang24ai.html · fetched 2026-08-29 · aec34a95f56b
- site_page: https://proceedings.mlr.press/faq.html · fetched 2026-08-29 · 8708020cd335
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| YangLing0818/RPG-DiffusionMaster | main | 28 |
For agents
markdown · JSON · MCP: product_card(name="YangLing0818/RPG-DiffusionMaster")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem