# character-ai/Ovi

Repository: https://github.com/character-ai/Ovi
Canonical: https://ross.abutalabs.com/products/ovi
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2025-11-15T02:55:20+00:00

## Health v2 (maintenance only)
Score: 40/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 52, release rhythm 35, longevity 24
- inputs: {"age_days": 343, "days_push": 291, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1748, forks 204 (observed 2026-08-28T04:05:30.971331+00:00)

## What it is
Ovi is a video-plus-audio generation model from Character AI that simultaneously generates synchronized video and audio from text or text+image inputs, using a twin backbone cross-modal fusion architecture. It ships as a Python library with pretrained checkpoints (including a 5B audio branch) and supports 5- or 10-second videos at 960x960 resolution with ComfyUI integration.

## Use cases
- generate videos with synchronized audio from a text prompt
- create talking videos from an image plus text description
- generate 10-second 960x960 videos with sound effects and speech
- run a veo-3-like open-source video+audio generator locally
- integrate video-audio generation into ComfyUI workflows
- generate videos in different aspect ratios like 9:16 or 16:9

## When to choose
- you need open-source joint video and audio generation rather than video-only models
- you want text-to-video or image-to-video with synchronized audio in one pass
- you need a self-hosted alternative to proprietary video generation APIs
- you want ComfyUI or Hugging Face integration for generative video workflows

## When to avoid
- you only need video generation without an audio track
- you lack a capable GPU for large diffusion model inference
- you need long-form videos beyond 10 seconds
- you need production-grade video editing rather than generation

## Facets
- artifact type: library
- maturity: active
- function: video-processing, audio-processing, machine-learning, deep-learning, llm-inference
- domain: artificial-intelligence, deep-learning, media
- platform: python
- tags: video-generation, audio-generation, text-to-video, image-to-video, cross-modal-fusion, diffusion-model, comfyui, tts, generative-ai, video, audio, gpu, linux, docker

## Member repositories
- character-ai/Ovi (main) score 40

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:30.971331+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:28:49.449640+00:00, confidence not recorded.
  - readme: https://github.com/character-ai/Ovi (fetched 2026-08-28T04:05:30.971331+00:00, sha e66da14ed9a7)
- Data as of 2026-08-30T08:39:29.467469+00:00.
