# OpenMOSS/MOVA

MOVA: Towards Scalable and Synchronized Video–Audio Generation

Repository: https://github.com/OpenMOSS/MOVA
Canonical: https://ross.abutalabs.com/products/mova
Homepage: https://mosi.cn/models/mova
Language: Python
License: Apache-2.0
License Family: permissive
Topics: video-audio-generation, diffusion-models, multimodal, sglang
Last push: 2026-08-31T06:43:10+00:00

## Health v2 (maintenance only)
Score: 60/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 100, release rhythm 35, longevity 15
- inputs: {"age_days": 217, "days_push": 2, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1106, forks 91 (observed 2026-09-01T02:13:57.461081+00:00)

## What it is
MOVA is an open-source foundation model and toolkit for joint video-audio generation, synthesizing synchronized video and audio in a single diffusion-based inference pass. It ships model weights, inference code, training and LoRA fine-tuning pipelines, evaluation code, a benchmark dataset, and ComfyUI integration.

## Use cases
- generate videos with synchronized audio from text prompts
- create lip-synced talking videos in multiple languages
- add environment-aware sound effects to generated video
- fine-tune video-audio generation with LoRA on custom data
- evaluate video-audio generation models on a benchmark
- generate video programmatically via API
- use video-audio generation inside ComfyUI workflows

## When to choose
- you need open-source video generation with native synchronized audio rather than a cascaded pipeline
- you want lip-sync and sound effects without closed models like Sora 2 or Veo 3
- you need weights, training code, and fine-tuning scripts to build on
- you want reproducible evaluation with a released benchmark

## When to avoid
- you only need video without audio and prefer lighter established text-to-video models
- you lack GPU hardware for large diffusion model inference
- you need a polished end-user product rather than a research toolkit

## Facets
- artifact type: library
- maturity: active
- function: video-processing, audio-processing, deep-learning, llm-inference, stable-diffusion
- domain: artificial-intelligence, deep-learning, machine-learning
- platform: python
- tags: video-audio-generation, diffusion-models, multimodal-generation, lip-sync, sound-effects, lora-fine-tuning, comfyui, sglang, foundation-model, text-to-video, video, audio, gpu, linux, docker

## Member repositories
- OpenMOSS/MOVA (main) score 60

## Provenance
- Observed fields: from GitHub, fetched 2026-09-01T02:13:57.461081+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:44:42.909473+00:00, confidence not recorded.
  - readme: https://github.com/OpenMOSS/MOVA (fetched 2026-09-01T02:13:57.461081+00:00, sha 0812e7584e51)
  - homepage: https://mosi.cn/models/mova (fetched 2026-08-29T12:48:17.744312+00:00, sha aaf084a1e7e0)
- Data as of 2026-08-30T08:39:29.467469+00:00.
