OpenMOSS/MOVA
MOVA: Towards Scalable and Synchronized Video–Audio Generation observed · 2026-09-01
Health v2 · maintenance only
60/100
- Activity 100
- Release rhythm 35
- Longevity 15
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 217
- days_rel: n/a
- days_push: 2
- n_releases_24m: 0
Adoption not part of the score
1106 stars · 91 forks observed · 2026-09-01
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
MOVA is an open-source foundation model and toolkit for joint video-audio generation, synthesizing synchronized video and audio in a single diffusion-based inference pass. It ships model weights, inference code, training and LoRA fine-tuning pipelines, evaluation code, a benchmark dataset, and ComfyUI integration.
Use cases
- generate videos with synchronized audio from text prompts
- create lip-synced talking videos in multiple languages
- add environment-aware sound effects to generated video
- fine-tune video-audio generation with LoRA on custom data
- evaluate video-audio generation models on a benchmark
- generate video programmatically via API
- use video-audio generation inside ComfyUI workflows
When to choose
- you need open-source video generation with native synchronized audio rather than a cascaded pipeline
- you want lip-sync and sound effects without closed models like Sora 2 or Veo 3
- you need weights, training code, and fine-tuning scripts to build on
- you want reproducible evaluation with a released benchmark
When to avoid
- you only need video without audio and prefer lighter established text-to-video models
- you lack GPU hardware for large diffusion model inference
- you need a polished end-user product rather than a research toolkit
Facets
library · maturity active
video-processing audio-processing deep-learning llm-inference stable-diffusion artificial-intelligence deep-learning machine-learning python video-audio-generation diffusion-models multimodal-generation lip-sync sound-effects lora-fine-tuning comfyui sglang foundation-model text-to-video video audio gpu linux docker
2 sources
- readme: https://github.com/OpenMOSS/MOVA · fetched 2026-09-01 · 0812e7584e51
- homepage: https://mosi.cn/models/mova · fetched 2026-08-29 · aaf084a1e7e0
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| OpenMOSS/MOVA | main | 60 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem