Ross ROSS = Recommend OSS · open-source software intelligence for agents

fishaudio/fish-speech

SOTA Open Source TTS observed · 2026-08-28

github.com/fishaudio/fish-speech · homepage · Python · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

74/100

  • Activity 99
  • Release rhythm 40
  • Longevity 75

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 29.5
  • age_days: 1058
  • days_rel: 459
  • days_push: 11
  • n_releases_24m: 7

Full methodology

Adoption not part of the score

32413 stars · 2797 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 languages, using a Dual-Autoregressive architecture with reinforcement learning alignment. It supports voice cloning, inline prosody/emotion control via natural-language tags, multi-speaker and multi-turn generation, with CLI, WebUI, and API server deployment options.

Use cases

  • generate natural speech from text
  • clone a voice from a short audio sample
  • build a multilingual TTS API server
  • add expressive emotional narration with tags like [laugh] and [whispers]
  • generate multi-speaker dialogue audio
  • self-host a text-to-speech service with Docker

When to choose

  • you need state-of-the-art TTS quality with low word error rates
  • you want fine-grained emotion and prosody control in generated speech
  • you need multilingual speech generation across ~50 languages
  • you can run inference on a 24GB GPU and want self-hosted deployment

When to avoid

  • you need a permissively licensed model for unrestricted commercial use (custom research license applies)
  • you only have CPU-only or low-VRAM hardware
  • you need real-time on-device TTS for mobile or embedded devices

Facets

library · maturity active

tts speech-recognition machine-learning llm-inference audio-processing speech-processing deep-learning artificial-intelligence python cross-platform voice-cloning text-to-speech zero-shot-tts multilingual emotion-control vqgan autoregressive audio linux docker gpu

3 sources

Member repositories

RepositoryRoleHealth v2
fishaudio/fish-speechmain74

For agents

markdown · JSON · MCP: product_card(name="fishaudio/fish-speech")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem