# MeiGen-AI/MultiTalk

[NeurIPS 2025] Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Repository: https://github.com/MeiGen-AI/MultiTalk
Canonical: https://ross.abutalabs.com/products/multitalk
Homepage: https://meigen-ai.github.io/multi-talk/
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2026-05-22T02:38:22+00:00

## Health v2 (maintenance only)
Score: 56/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 83, release rhythm 35, longevity 33
- inputs: {"age_days": 462, "days_push": 104, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2992, forks 497 (observed 2026-08-28T04:07:35.717806+00:00)

## What it is
MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a text prompt, with lip motions synchronized to each speaker. It supports conversation, singing, interaction control, and cartoon-style video generation.

## Use cases
- generate talking head videos from audio
- create multi-person conversation videos with lip sync
- make a person sing from an audio track
- animate a reference photo to speak given dialogue audio
- generate cartoon character conversation videos
- control interactions between people in generated video

## When to choose
- you need audio-driven video generation with accurate per-speaker lip sync
- you want to animate multiple people in one video from separate audio streams
- you need a research-grade, actively maintained model with published results (NeurIPS 2025)

## When to avoid
- you need real-time or low-latency video generation
- you lack a GPU or cannot run heavy diffusion models
- you only need single-image animation without audio alignment

## Facets
- artifact type: library
- maturity: active
- function: video-processing, machine-learning, deep-learning, audio-processing, image-processing
- domain: deep-learning, computer-vision, artificial-intelligence
- platform: python
- tags: talking-head-generation, lip-sync, audio-driven-video, multi-person-conversation, video-generation, diffusion-models, research-code, neurips-2025, video, audio, gpu, linux

## Member repositories
- MeiGen-AI/MultiTalk (main) score 56

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:35.717806+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:31:19.282811+00:00, confidence not recorded.
  - readme: https://github.com/MeiGen-AI/MultiTalk (fetched 2026-08-28T04:07:35.717806+00:00, sha a6ca070875ea)
  - homepage: https://meigen-ai.github.io/multi-talk/ (fetched 2026-08-29T09:46:09.541405+00:00, sha ca5dc3004980)
- Data as of 2026-08-30T08:39:29.467469+00:00.
