# CyberAgentAILab/TANGO

[ICLR 2025 Oral] TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio-Motion Embedding and Diffusion Interpolation

Repository: https://github.com/CyberAgentAILab/TANGO
Canonical: https://ross.abutalabs.com/products/cyberagentailab-tango
Homepage: https://pantomatrix.github.io/TANGO/
Language: Python
License: NOASSERTION
License Family: other
Last push: 2025-08-24T15:23:17+00:00

## Health v2 (maintenance only)
Score: 39/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 38, release rhythm 35, longevity 48
- inputs: {"age_days": 674, "days_push": 374, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1163, forks 149 (observed 2026-08-28T04:03:49.738302+00:00)

## What it is
TANGO is a research library from CyberAgent AI Lab that generates co-speech gesture videos by reenactment, using hierarchical audio-motion embeddings (AuMoClip) and diffusion interpolation. It was published as an ICLR 2025 Oral paper and includes inference and training code with pretrained weights.

## Use cases
- generate talking gesture videos from speech audio
- reenact a reference speaker video with new speech
- synthesize realistic body gestures synced to audio
- research on audio-driven human motion generation
- create virtual presenter videos from a script
- train audio-to-motion diffusion models

## When to choose
- you need audio-driven gesture video generation with pretrained weights
- you want to reproduce or extend an ICLR 2025 research result
- you have a GPU environment and Python 3.10/CUDA 11.8 available
- you want a Hugging Face demo-backed gesture synthesis pipeline

## When to avoid
- you need production-grade, licensed software (license is non-standard)
- you lack a CUDA GPU or cannot handle heavy model dependencies
- you need real-time or low-latency generation (inference takes minutes per clip)
- you want a polished end-user application rather than research code

## Facets
- artifact type: library
- maturity: active
- function: deep-learning, video-processing, audio-processing, machine-learning, llm-inference
- domain: deep-learning, computer-vision, artificial-intelligence
- platform: python
- tags: co-speech-gesture, video-reenactment, diffusion-models, audio-motion-embedding, talking-head, iclr-2025, gesture-generation, research-code, audio, video, linux, gpu, docker

## Member repositories
- CyberAgentAILab/TANGO (main) score 39

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:49.738302+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:31:32.384241+00:00, confidence not recorded.
  - readme: https://github.com/CyberAgentAILab/TANGO (fetched 2026-08-28T04:03:49.738302+00:00, sha 8fef059d015b)
  - homepage: https://pantomatrix.github.io/TANGO/ (fetched 2026-08-29T12:35:43.180151+00:00, sha f3c6a572446d)
- Data as of 2026-08-30T08:39:29.467469+00:00.
