# EchoMimic

[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

Repository: https://github.com/antgroup/echomimic_v3
Canonical: https://ross.abutalabs.com/products/echomimic
Homepage: https://antgroup.github.io/ai/echomimic_v3/
Language: Python
License: Apache-2.0
License Family: permissive
Topics: audio-driven-body-animation, audio-driven-portrait-animations, human-animation, video-generation
Last push: 2026-03-18T03:31:56+00:00
Link (homepage): https://antgroup.github.io/ai/echomimic_v3/

## Health v2 (maintenance only)
Score: 50/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 72, release rhythm 35, longevity 29
- inputs: {"age_days": 406, "days_push": 168, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1027, forks 122 (observed 2026-08-28T04:03:17.206950+00:00)

## What it is
EchoMimic is a series of open-source models (V1-V3) from Ant Group for audio-driven human animation, generating lifelike talking-head, portrait, and semi-body videos from a reference image and audio. EchoMimicV3 unifies multi-modal and multi-task human animation in a compact 1.3B-parameter model.

## Use cases
- generate a talking head video from a photo and audio clip
- animate a portrait image to lip-sync with speech
- create audio-driven semi-body human animation videos
- make an avatar sing from an audio track
- generate cartoon character animation driven by audio
- run human animation research experiments

## When to choose
- you need audio-driven talking-face or portrait animation with published research backing
- you want a compact unified model (V3) handling multiple animation tasks
- you need a permissively licensed (Apache-2.0) animation model with pretrained weights on HuggingFace

## When to avoid
- you need real-time animation on CPU-only hardware
- you need full-body animation rather than portrait or semi-body
- you want a polished end-user app rather than a research codebase requiring GPU setup

## Facets
- artifact type: library
- maturity: active
- function: video-processing, machine-learning, deep-learning, audio-processing, image-processing
- domain: artificial-intelligence, computer-vision, deep-learning
- platform: python, cross-platform
- tags: talking-head-generation, audio-driven-animation, human-animation, diffusion-models, cvpr2025, video-generation, portrait-animation, video, audio, gpu, linux

## Member repositories
- antgroup/echomimic_v3 (main) score 50
- antgroup/echomimic_v2 (mirror) score 52
- antgroup/echomimic (mirror) score 58

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:17.206950+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:19:29.648706+00:00, confidence not recorded.
  - readme: https://github.com/antgroup/echomimic_v3 (fetched 2026-08-28T04:03:17.206950+00:00, sha cd754d696ebe)
  - homepage: https://antgroup.github.io/ai/echomimic_v3/ (fetched 2026-08-29T09:04:01.167410+00:00, sha 7d2284611b55)
- Data as of 2026-08-30T08:39:29.467469+00:00.
