Ross ROSS = Recommend OSS · open-source software intelligence for agents

k2-fsa/OmniVoice

High-Quality Voice Cloning TTS for 600+ Languages observed · 2026-08-28

github.com/k2-fsa/OmniVoice · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

79/100

  • Activity 99
  • Release rhythm 93
  • Longevity 11

Flags: young

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 9
  • age_days: 155
  • days_rel: 48
  • days_push: 9
  • n_releases_24m: 6

Full methodology

Adoption not part of the score

9455 stars · 1551 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style architecture. It provides high-quality voice cloning, voice design via speaker attributes, and fast inference (up to 40x real-time) through Python APIs and command-line tools.

Use cases

  • clone a voice from a short audio sample
  • generate speech in hundreds of languages from text
  • design a custom voice by specifying gender, age, pitch, and accent
  • synthesize speech with fine-grained control like [laughter] tags and phoneme pronunciation correction
  • run fast TTS inference on GPU or Apple Silicon
  • fine-tune or evaluate a multilingual TTS model

When to choose

  • you need zero-shot TTS with the broadest language coverage (600+ languages)
  • you want state-of-the-art voice cloning quality
  • you need very fast inference (RTF as low as 0.025)
  • you want fine-grained control over non-verbal sounds and pronunciation
  • you prefer an Apache-2.0 licensed model with Python API and CLI

When to avoid

  • you need a lightweight CPU-only TTS without GPU acceleration
  • you need a simple single-language TTS with minimal setup
  • you want a hosted SaaS API rather than running models locally
  • you need real-time streaming synthesis specifically

Facets

library · maturity active

tts speech-recognition machine-learning deep-learning audio-processing speech-processing machine-learning artificial-intelligence python cross-platform cli voice-cloning zero-shot-tts multilingual diffusion-language-model voice-design speech-synthesis natural-language-processing gpu

2 sources

Member repositories

RepositoryRoleHealth v2
k2-fsa/OmniVoicemain79

For agents

markdown · JSON · MCP: product_card(name="k2-fsa/OmniVoice")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem