# chengzeyi/stable-fast

https://wavespeed.ai/ Best inference performance optimization framework for HuggingFace Diffusers on NVIDIA GPUs.

Repository: https://github.com/chengzeyi/stable-fast
Canonical: https://ross.abutalabs.com/products/stable-fast
Language: Python
License: MIT
License Family: permissive
Topics: cuda, diffusers, pytorch, stable-diffusion, deeplearnng, inference-engines, openai-triton, performance-optimizations, torch, stable-video-diffusion
Last push: 2025-03-27T08:07:20+00:00

## Health v2 (maintenance only)
Score: 24/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 13, release rhythm 8, longevity 75
- inputs: {"age_days": 1051, "days_push": 524, "days_rel": 644, "gap_med": null, "n_releases_24m": 1}
- flags: prerelease_only
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1302, forks 93 (observed 2026-08-28T04:04:18.266147+00:00)

## What it is
Stable Fast is an ultra lightweight performance optimization framework for HuggingFace Diffusers pipelines on NVIDIA GPUs. It achieves SOTA inference speed for diffusion models with only seconds of compilation, supporting dynamic shapes, LoRA, and ControlNet out of the box.

## Use cases
- speed up stable diffusion image generation on nvidia gpus
- optimize huggingface diffusers pipelines for faster inference
- run flux or stable video diffusion with low latency
- avoid long tensorrt compile times for diffusion models
- accelerate diffusers with lora and controlnet support
- serve stable diffusion with minimal compilation overhead

## When to choose
- you need fast diffusion model inference on NVIDIA CUDA GPUs
- you want seconds-level compilation instead of TensorRT's minutes
- you need dynamic shapes, LoRA, or ControlNet with optimized pipelines

## When to avoid
- you need non-CUDA hardware backends
- you want actively developed support for newest models like SD3 or Sora-like architectures
- you need a maintained project - development is paused in favor of newer torch._dynamo-based work

## Facets
- artifact type: library
- maturity: maintenance
- function: llm-inference, machine-learning, gpu-computing, deep-learning
- domain: machine-learning, deep-learning, image-processing, gpu-computing, performance
- platform: python
- tags: stable-diffusion, diffusers, inference-optimization, cuda, pytorch, triton, image-generation, video-generation, gpu, linux, docker

## Member repositories
- chengzeyi/stable-fast (main) score 24

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:18.266147+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:51:16.806818+00:00, confidence not recorded.
  - readme: https://github.com/chengzeyi/stable-fast (fetched 2026-08-28T04:04:18.266147+00:00, sha dbc85c35cbd7)
  - registry_pypi: https://pypi.org/pypi/stable-fast/json (fetched 2026-08-29T12:09:29.138545+00:00, sha e3861cace93c)
- Data as of 2026-08-30T08:39:29.467469+00:00.
