Ross ROSS = Recommend OSS · open-source software intelligence for agents

Plachtaa/VALL-E-X

An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/ observed · 2026-08-28

github.com/Plachtaa/VALL-E-X · Python · MIT (permissive) · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 0
  • Release rhythm 35
  • Longevity 80

Flags: no_releases archived

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1131
  • days_rel: n/a
  • days_push: 934
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

7931 stars · 781 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint. It supports multilingual speech synthesis, voice cloning from short prompts, and emotional speech generation.

Use cases

  • clone a voice from a short audio sample
  • generate speech from text in English, Chinese, or Japanese
  • synthesize emotional speech
  • run zero-shot text-to-speech locally on GPU
  • create multilingual voiceovers

When to choose

  • you need zero-shot TTS with voice cloning from a few seconds of audio
  • you want a free, self-hosted alternative to commercial TTS APIs
  • you need multilingual synthesis with cross-lingual voice transfer

When to avoid

  • you need production-grade, actively maintained TTS with support
  • you have no GPU, since inference requires CUDA
  • you need languages beyond English, Chinese, and Japanese

Facets

library · maturity maintenance

tts speech-recognition deep-learning machine-learning speech-processing artificial-intelligence python cross-platform voice-cloning zero-shot-tts text-to-speech emotional-speech transformer multilingual pretrained-model audio gpu

1 source

Member repositories

RepositoryRoleHealth v2
Plachtaa/VALL-E-Xmain10

For agents

markdown · JSON · MCP: product_card(name="Plachtaa/VALL-E-X")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem