Ross ROSS = Recommend OSS · open-source software intelligence for agents

WhisperSpeech/WhisperSpeech

An Open Source text-to-speech system built by inverting Whisper. observed · 2026-08-28

github.com/WhisperSpeech/WhisperSpeech · homepage · Jupyter Notebook · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

56/100

  • Activity 57
  • Release rhythm 35
  • Longevity 92

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1296
  • days_rel: n/a
  • days_push: 262
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

4639 stars · 275 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

WhisperSpeech is an open-source text-to-speech system built by inverting OpenAI's Whisper model, aiming to be 'Stable Diffusion for speech'. It supports voice cloning, multilingual synthesis, and runs faster than real time on consumer GPUs, with all code and training data properly licensed for commercial use.

Use cases

  • generate speech from text with an open-source model
  • clone a voice from a short audio sample
  • synthesize multilingual or code-switching sentences
  • run fast TTS locally on a consumer GPU
  • build a commercially safe TTS product without licensing risk
  • fine-tune or hack a speech synthesis pipeline

When to choose

  • you need open-source, commercially usable TTS with permissive licenses
  • you want voice cloning and multilingual support
  • you have a GPU and want faster-than-real-time local synthesis
  • you want a hackable research-friendly speech model

When to avoid

  • you need production TTS with many polished prebuilt voices out of the box
  • you have no GPU and need low-latency CPU inference
  • you need broad language coverage today (English is the main current release)

Facets

library · maturity active

tts machine-learning deep-learning audio-processing speech-processing artificial-intelligence deep-learning python cross-platform speech-synthesis voice-cloning pytorch whisper open-source-models audio-generation audio gpu

2 sources

Member repositories

RepositoryRoleHealth v2
WhisperSpeech/WhisperSpeechmain56

For agents

markdown · JSON · MCP: product_card(name="WhisperSpeech/WhisperSpeech")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem