Ross ROSS = Recommend OSS · open-source software intelligence for agents

chengzeyi/stable-fast

https://wavespeed.ai/ Best inference performance optimization framework for HuggingFace Diffusers on NVIDIA GPUs. observed · 2026-08-28

github.com/chengzeyi/stable-fast · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

24/100

  • Activity 13
  • Release rhythm 8
  • Longevity 75

Flags: prerelease_only

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1051
  • days_rel: 644
  • days_push: 524
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

1302 stars · 93 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Stable Fast is an ultra lightweight performance optimization framework for HuggingFace Diffusers pipelines on NVIDIA GPUs. It achieves SOTA inference speed for diffusion models with only seconds of compilation, supporting dynamic shapes, LoRA, and ControlNet out of the box.

Use cases

  • speed up stable diffusion image generation on nvidia gpus
  • optimize huggingface diffusers pipelines for faster inference
  • run flux or stable video diffusion with low latency
  • avoid long tensorrt compile times for diffusion models
  • accelerate diffusers with lora and controlnet support
  • serve stable diffusion with minimal compilation overhead

When to choose

  • you need fast diffusion model inference on NVIDIA CUDA GPUs
  • you want seconds-level compilation instead of TensorRT's minutes
  • you need dynamic shapes, LoRA, or ControlNet with optimized pipelines

When to avoid

  • you need non-CUDA hardware backends
  • you want actively developed support for newest models like SD3 or Sora-like architectures
  • you need a maintained project - development is paused in favor of newer torch._dynamo-based work

Facets

library · maturity maintenance

llm-inference machine-learning gpu-computing deep-learning machine-learning deep-learning image-processing gpu-computing performance python stable-diffusion diffusers inference-optimization cuda pytorch triton image-generation video-generation gpu linux docker

2 sources

Member repositories

RepositoryRoleHealth v2
chengzeyi/stable-fastmain24

For agents

markdown · JSON · MCP: product_card(name="chengzeyi/stable-fast")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem