# facebookresearch/audio2photoreal

Code and dataset for photorealistic Codec Avatars driven from audio

Repository: https://github.com/facebookresearch/audio2photoreal
Canonical: https://ross.abutalabs.com/products/audio2photoreal
Language: Python
License: NOASSERTION
License Family: other
Archived: true
Last push: 2024-09-15T02:29:57+00:00

## Health v2 (maintenance only)
Score: 10/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 0, release rhythm 8, longevity 69
- inputs: {"age_days": 974, "days_push": 718, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: archived, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2850, forks 282 (observed 2026-08-28T04:07:25.634830+00:00)

## What it is
A PyTorch codebase and dataset from Meta Research for generating photorealistic conversational avatars (face and body motion) driven from audio. It includes training and inference code, pretrained diffusion and VQ-VAE models, and a Gradio demo for rendering avatars from recorded audio.

## Use cases
- generate a photorealistic avatar video from an audio clip
- train audio-driven body and face motion models
- research conversational avatar gesture synthesis
- visualize human conversation motion capture data
- run a demo that renders talking humans from microphone input
- reproduce CVPR paper results on audio-to-embodiment synthesis

## When to choose
- you need research-grade audio-driven avatar motion generation with pretrained models
- you want the accompanying conversational gesture dataset for training
- you want to extend or reproduce state-of-the-art photorealistic avatar research

## When to avoid
- you need a production-ready avatar SDK or end-user product
- you lack a CUDA GPU or the rendering pipeline dependencies
- you need real-time low-latency avatar animation in a commercial app
- you want simple text-to-avatar generation without audio

## Facets
- artifact type: dataset
- maturity: active
- function: machine-learning, deep-learning, audio-processing, computer-vision, graphics, simulation
- domain: machine-learning, computer-vision, graphics, artificial-intelligence
- platform: python
- tags: codec-avatars, audio-driven-animation, diffusion-models, human-motion, conversational-agents, pytorch, research-code, face-and-body-generation, gpu, linux

## Member repositories
- facebookresearch/audio2photoreal (main) score 10

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:25.634830+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:36:59.812461+00:00, confidence not recorded.
  - readme: https://github.com/facebookresearch/audio2photoreal (fetched 2026-08-28T04:07:25.634830+00:00, sha 2249b55320af)
- Data as of 2026-08-30T08:39:29.467469+00:00.
