QwenAudio/CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability. observed · 2026-08-28
Health v2 · maintenance only
61/100
- Activity 84
- Release rhythm 35
- Longevity 56
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 791
- days_rel: n/a
- days_push: 100
- n_releases_24m: 0
Adoption not part of the score
22925 stars · 2637 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-shot voice cloning, streaming synthesis with ~150ms latency, and instruction-based control across 9 languages and 18+ Chinese dialects.
Use cases
- clone a voice from a short audio sample
- generate speech in Chinese, English, Japanese, Korean and other languages
- build a streaming text-to-speech service with low latency
- fine-tune a TTS model on custom speaker data
- control emotion, speed, and dialect of synthesized speech via instructions
- synthesize speech with correct pronunciation of numbers and symbols
When to choose
- you need production-grade multilingual or cross-lingual TTS with voice cloning
- you want streaming audio output with very low latency
- you need Chinese dialect support or pronunciation inpainting
- you want to train or fine-tune your own voice generation model
When to avoid
- you need a lightweight offline TTS without GPU resources
- you only need simple pre-recorded voice playback
- you require languages outside its supported set
- you want a fully managed cloud TTS API rather than self-hosted models
Facets
library · maturity active
tts speech-recognition machine-learning llm-inference audio-processing speech-processing artificial-intelligence large-language-models python cross-platform voice-cloning text-to-speech multilingual zero-shot-synthesis streaming-tts fine-tuning voice-generation chinese-dialects natural-language-processing audio gpu linux
1 source
- readme: https://github.com/QwenAudio/CosyVoice · fetched 2026-08-28 · 5e932c1da217
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| QwenAudio/CosyVoice | main | 61 |
For agents
markdown · JSON · MCP: product_card(name="QwenAudio/CosyVoice")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem