Ross ROSS = Recommend OSS · open-source software intelligence for agents

character-ai/Ovi

None observed · 2026-08-28

github.com/character-ai/Ovi · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

40/100

  • Activity 52
  • Release rhythm 35
  • Longevity 24

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 343
  • days_rel: n/a
  • days_push: 291
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1748 stars · 204 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Ovi is a video-plus-audio generation model from Character AI that simultaneously generates synchronized video and audio from text or text+image inputs, using a twin backbone cross-modal fusion architecture. It ships as a Python library with pretrained checkpoints (including a 5B audio branch) and supports 5- or 10-second videos at 960x960 resolution with ComfyUI integration.

Use cases

  • generate videos with synchronized audio from a text prompt
  • create talking videos from an image plus text description
  • generate 10-second 960x960 videos with sound effects and speech
  • run a veo-3-like open-source video+audio generator locally
  • integrate video-audio generation into ComfyUI workflows
  • generate videos in different aspect ratios like 9:16 or 16:9

When to choose

  • you need open-source joint video and audio generation rather than video-only models
  • you want text-to-video or image-to-video with synchronized audio in one pass
  • you need a self-hosted alternative to proprietary video generation APIs
  • you want ComfyUI or Hugging Face integration for generative video workflows

When to avoid

  • you only need video generation without an audio track
  • you lack a capable GPU for large diffusion model inference
  • you need long-form videos beyond 10 seconds
  • you need production-grade video editing rather than generation

Facets

library · maturity active

video-processing audio-processing machine-learning deep-learning llm-inference artificial-intelligence deep-learning media python video-generation audio-generation text-to-video image-to-video cross-modal-fusion diffusion-model comfyui tts generative-ai video audio gpu linux docker

1 source

Member repositories

RepositoryRoleHealth v2
character-ai/Ovimain40

For agents

markdown · JSON · MCP: product_card(name="character-ai/Ovi")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem