Ross ROSS = Recommend OSS · open-source software intelligence for agents

CyberAgentAILab/TANGO

[ICLR 2025 Oral] TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio-Motion Embedding and Diffusion Interpolation observed · 2026-08-28

github.com/CyberAgentAILab/TANGO · homepage · Python · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

39/100

  • Activity 38
  • Release rhythm 35
  • Longevity 48

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 674
  • days_rel: n/a
  • days_push: 374
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1163 stars · 149 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

TANGO is a research library from CyberAgent AI Lab that generates co-speech gesture videos by reenactment, using hierarchical audio-motion embeddings (AuMoClip) and diffusion interpolation. It was published as an ICLR 2025 Oral paper and includes inference and training code with pretrained weights.

Use cases

  • generate talking gesture videos from speech audio
  • reenact a reference speaker video with new speech
  • synthesize realistic body gestures synced to audio
  • research on audio-driven human motion generation
  • create virtual presenter videos from a script
  • train audio-to-motion diffusion models

When to choose

  • you need audio-driven gesture video generation with pretrained weights
  • you want to reproduce or extend an ICLR 2025 research result
  • you have a GPU environment and Python 3.10/CUDA 11.8 available
  • you want a Hugging Face demo-backed gesture synthesis pipeline

When to avoid

  • you need production-grade, licensed software (license is non-standard)
  • you lack a CUDA GPU or cannot handle heavy model dependencies
  • you need real-time or low-latency generation (inference takes minutes per clip)
  • you want a polished end-user application rather than research code

Facets

library · maturity active

deep-learning video-processing audio-processing machine-learning llm-inference deep-learning computer-vision artificial-intelligence python co-speech-gesture video-reenactment diffusion-models audio-motion-embedding talking-head iclr-2025 gesture-generation research-code audio video linux gpu docker

2 sources

Member repositories

RepositoryRoleHealth v2
CyberAgentAILab/TANGOmain39

For agents

markdown · JSON · MCP: product_card(name="CyberAgentAILab/TANGO")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem