# Francis-Rings/StableAvatar

We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-processing, conditioned on a reference image and audio.

Repository: https://github.com/Francis-Rings/StableAvatar
Canonical: https://ross.abutalabs.com/products/stableavatar
Language: Python
License: MIT
License Family: permissive
Topics: aigc, avatar-generator, video-generation
Last push: 2026-01-20T01:44:41+00:00

## Health v2 (maintenance only)
Score: 46/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 63, release rhythm 35, longevity 27
- inputs: {"age_days": 387, "days_push": 226, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1258, forks 113 (observed 2026-08-28T04:04:09.492938+00:00)

## What it is
StableAvatar is an end-to-end video diffusion transformer that generates infinite-length, high-quality talking avatar videos from a reference image and audio, with no post-processing. It is a Python research codebase with pretrained models on Hugging Face and an online demo.

## Use cases
- generate a talking avatar video from a photo and audio clip
- create infinite-length lip-synced avatar videos
- audio-driven digital human video generation
- make an AI presenter video from one reference image
- research on audio-conditioned video diffusion models

## When to choose
- you need long, identity-consistent talking-head videos from a single image plus audio
- you want a research-grade open model for audio-driven avatar generation
- you have a GPU environment and want to run inference locally

## When to avoid
- you need real-time avatar animation or live streaming
- you lack a capable GPU or want a lightweight CPU tool
- you need fine-grained manual animation control rather than audio-driven synthesis

## Facets
- artifact type: library
- maturity: active
- function: video-processing, machine-learning, deep-learning, tts
- domain: artificial-intelligence, deep-learning, image-processing
- platform: python
- tags: video-diffusion, avatar-generation, talking-head, audio-driven, diffusion-transformer, aigc, video, gpu, linux

## Member repositories
- Francis-Rings/StableAvatar (main) score 46

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:09.492938+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T05:05:33.571130+00:00, confidence not recorded.
  - readme: https://github.com/Francis-Rings/StableAvatar (fetched 2026-08-28T04:04:09.492938+00:00, sha cff3d6d1622f)
- Data as of 2026-08-30T08:39:29.467469+00:00.
