# lenML/Speech-AI-Forge

🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.

Repository: https://github.com/lenML/Speech-AI-Forge
Canonical: https://ross.abutalabs.com/products/speech-ai-forge
Language: Python
License: AGPL-3.0
License Family: copyleft
Topics: chattts, ssml, tts, chattts-forge, agent, gpt, llm, text-to-speech, colab, llama, chinese, english, fish-speech, cosyvoice, cosy-voice, asr, stt, firered, whisper, fireredtts
Last push: 2026-05-21T12:49:07+00:00

## Health v2 (maintenance only)
Score: 62/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 83, release rhythm 36, longevity 58
- inputs: {"age_days": 823, "days_push": 104, "days_rel": 212, "gap_med": null, "n_releases_24m": 1}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1416, forks 189 (observed 2026-08-28T04:04:39.874827+00:00)

## What it is
Speech-AI-Forge is a Python project built around multiple TTS generation models (ChatTTS, CosyVoice, Fish-Speech, Index-TTS, F5-TTS, FireRedTTS, and more) that provides both an API server and a Gradio-based WebUI. It also supports ASR/STT via Whisper and SenseVoice, SSML, and cloud TTS backends like MiniMax.

## Use cases
- generate speech from text with multiple tts models
- self-host a text-to-speech api server
- clone a voice for tts synthesis
- transcribe audio to text with whisper or sensevoice
- run a gradio webui for text-to-speech
- use ssml to control speech synthesis
- try tts models in colab without installing

## When to choose
- you want one self-hosted server exposing many open-source TTS models behind a unified API
- you need both a WebUI for experimentation and an API for integration
- you want SSML support and voice cloning across multiple TTS engines
- you need Chinese and English speech synthesis

## When to avoid
- you need a lightweight production TTS service with a single optimized model
- you require a permissively licensed dependency since it is AGPL-3.0
- you need real-time low-latency streaming TTS at scale
- you only want cloud TTS without local model support

## Facets
- artifact type: application
- maturity: active
- function: tts, speech-recognition, http-server, llm-inference, api-framework
- domain: speech-processing, artificial-intelligence
- platform: python, windows, cross-platform
- tags: text-to-speech, gradio-webui, chattts, cosyvoice, fish-speech, ssml, voice-cloning, asr, whisper, colab, natural-language-processing, audio, docker, web-server, gpu

## Member repositories
- lenML/Speech-AI-Forge (main) score 62

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:39.874827+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:38:05.944629+00:00, confidence not recorded.
  - readme: https://github.com/lenML/Speech-AI-Forge (fetched 2026-08-28T04:04:39.874827+00:00, sha 74959f2b94d9)
- Data as of 2026-08-30T08:39:29.467469+00:00.
