Ross ROSS = Recommend OSS · open-source software intelligence for agents

YangLing0818/RPG-DiffusionMaster

[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG) observed · 2026-08-28

github.com/YangLing0818/RPG-DiffusionMaster · homepage · Jupyter Notebook · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

28/100

  • Activity 4
  • Release rhythm 35
  • Longevity 68

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 955
  • days_rel: n/a
  • days_push: 578
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1842 stars · 100 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Official implementation of RPG (Recaption, Plan, Generate), a training-free framework that uses multimodal LLMs as prompt recaptioners and region planners to improve compositional text-to-image generation and editing with diffusion models. It supports proprietary and open-source MLLMs and arbitrary diffusion backbones, enabling high-resolution, multi-object image generation.

Use cases

  • generate complex multi-object images from long text prompts
  • improve compositional text-to-image generation with diffusion models
  • edit images with text-guided closed-loop generation
  • generate super high resolution images with regional diffusion
  • use GPT-4 or local MLLMs to plan image subregions
  • research training-free prompt planning for diffusion models

When to choose

  • you need better multi-object composition and text-image alignment than vanilla Stable Diffusion or SDXL
  • you want a training-free framework compatible with various MLLMs and diffusion backbones
  • you need high-resolution or region-wise controlled image generation and editing

When to avoid

  • you need a production-ready end-user image generation app rather than research code
  • you cannot access an MLLM API or run a local multimodal model
  • you need fast single-prompt generation without LLM planning overhead

Facets

library · maturity active

llm-inference image-processing prompt-engineering machine-learning artificial-intelligence image-processing large-language-models deep-learning python cross-platform text-to-image diffusion-models multimodal-llm stable-diffusion image-editing research-code icml-2024 training-free gpu

3 sources

Member repositories

RepositoryRoleHealth v2
YangLing0818/RPG-DiffusionMastermain28

For agents

markdown · JSON · MCP: product_card(name="YangLing0818/RPG-DiffusionMaster")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem