Ross ROSS = Recommend OSS · open-source software intelligence for agents

vibevoice-community/VibeVoice

VibeVoice: Expressive, longform conversational speech synthesis. (Community fork) observed · 2026-08-28

github.com/vibevoice-community/VibeVoice · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

60/100

  • Activity 96
  • Release rhythm 35
  • Longevity 25

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 363
  • days_rel: n/a
  • days_push: 26
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1542 stars · 704 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker audio such as podcasts for up to 90 minutes with up to 4 speakers. It uses continuous speech tokenizers and a next-token diffusion LLM framework, with community-added fine-tuning and voice cloning support.

Use cases

  • generate a podcast from a multi-speaker script
  • synthesize long-form conversational audio with consistent voices
  • fine-tune a TTS model on a new language or voice
  • clone a voice for speech synthesis
  • run real-time streaming text-to-speech
  • create audiobook narration with multiple characters

When to choose

  • you need expressive multi-speaker conversational TTS beyond a few minutes
  • you want to fine-tune or voice-clone an open TTS model
  • you need podcast-style audio generation from text

When to avoid

  • you need lightweight on-device TTS without a GPU
  • you need only short single-phrase synthesis
  • you require official vendor support rather than a community fork

Facets

library · maturity active

tts speech-recognition machine-learning llm-inference audio-processing speech-processing artificial-intelligence deep-learning large-language-models python cross-platform text-to-speech voice-cloning podcast-generation multi-speaker long-form-audio fine-tuning diffusion community-fork audio gpu

2 sources

Member repositories

RepositoryRoleHealth v2
vibevoice-community/VibeVoicemain60

For agents

markdown · JSON · MCP: product_card(name="vibevoice-community/VibeVoice")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem