Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: tts

498 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
harry0703/MoneyPrinterTurbo
MoneyPrinterTurbo is an AI-powered application that generates high-definition short videos from a topic or keyword. It automatically writes…
92116915active
RVC-Boss/GPT-SoVITS
GPT-SoVITS is a Python-based few-shot voice cloning and text-to-speech system with an integrated WebUI. It supports zero-shot TTS from a 5-…
6861255active
microsoft/VibeVoice
VibeVoice is Microsoft's open-source frontier voice AI framework combining a next-token diffusion text-to-speech model for expressive, long…
6053225active
jamiepine/voicebox
Voicebox is a free, open-source, local-first AI voice studio desktop app that combines voice cloning, text-to-speech across 7 engines, and …
7551551active
moeru-ai/airi
Project AIRI is an open-source, self-hosted AI companion (waifu/virtual character) platform inspired by Neuro-sama, supporting Live2D and V…
8648451active
2noise/ChatTTS
ChatTTS is a generative text-to-speech model optimized for conversational dialogue scenarios such as LLM assistants, supporting English and…
6939797active
myshell-ai/OpenVoice
OpenVoice is a Python library and audio foundation model for instant voice cloning, requiring only a short reference audio clip to replicat…
3437312active
OpenBMB/VoxCPM
VoxCPM is a tokenizer-free text-to-speech system built on a diffusion autoregressive architecture that generates continuous speech represen…
7936145active
SillyTavern/SillyTavern
SillyTavern is a locally installed, self-hosted web UI for interacting with text generation LLMs, image generation engines, and TTS voice m…
8732691active
fishaudio/fish-speech
Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 l…
7432413active
koodo-reader/koodo-reader
Koodo Reader is a cross-platform ebook manager and reader supporting EPUB, PDF, Kindle, comic archives, and many other formats. It offers c…
9427983active
ATH-MaaS/Pixelle-Video
Pixelle-Video is an AI-powered fully automated short video engine that turns a single topic into a finished video. It automatically writes …
6727344active
resemble-ai/chatterbox
Chatterbox is a family of state-of-the-art open-source text-to-speech models by Resemble AI, including multilingual and low-latency Turbo v…
5226163active
readest/readest
Readest is an open-source, cross-platform ebook reader built with Next.js and Tauri v2, supporting EPUB, PDF, MOBI, AZW3, FB2, CBZ, and mor…
8423766active
index-tts/index-tts
IndexTTS is an industrial-level zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest IndexTTS…
7823511active
QwenAudio/CosyVoice
CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-…
6122925active
DrewThomasson/ebook2audiobook
A Python application that converts non-DRM e-books (e.g., EPUB, PDF) into audiobooks with chapters and metadata using TTS engines like XTTS…
9320053active
nari-labs/dia
Dia is a 1.6B-parameter open-weight text-to-speech model from Nari Labs that generates ultra-realistic multi-speaker dialogue in a single p…
4319378active
jianchang512/pyvideotrans
pyVideoTrans is an open-source desktop application (with WebUI and CLI modes) that translates videos from one language to another. It provi…
9018812active
NVIDIA-NeMo/Speech
NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, …
9818337active
Huanshere/VideoLingo
VideoLingo is an all-in-one AI video translation, localization, and dubbing tool that produces Netflix-quality single-line subtitles with w…
8018266active
leon-ai/leon
Leon is an open-source personal AI assistant that runs on your own server, combining tools, memory, and agentic execution with speech-to-te…
8617462active
KittenML/KittenTTS
KittenTTS is an open-source, ultra-lightweight text-to-speech library built on ONNX, with models ranging from 15M to 80M parameters (25-80 …
6615403active
SWivid/F5-TTS
F5-TTS is the official implementation of a fully non-autoregressive text-to-speech system based on flow matching with a Diffusion Transform…
8515167active
duixcom/Duix-Avatar
Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea…
5714871active
SesameAILabs/csm
CSM (Conversational Speech Model) is Sesame's speech generation model that produces conversational audio from text and audio context, using…
3014720active
k2-fsa/sherpa-onnx
sherpa-onnx is an offline speech processing toolkit built on next-gen Kaldi and onnxruntime, supporting speech-to-text, text-to-speech, spe…
9514411active
tisfeng/Easydict
Easydict is a concise and elegant macOS dictionary and translation app for looking up words and translating text, ready to use out of the b…
9514372stable
Open-LLM-VTuber/Open-LLM-VTuber
Open-LLM-VTuber is a cross-platform AI companion application that enables real-time voice conversations with LLMs, featuring a Live2D avata…
6713473active
livekit/agents
LiveKit Agents is an open-source framework for building realtime, multimodal voice AI agents that run as programmable participants in LiveK…
9013176active
QwenLM/Qwen3-TTS
Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,…
4813113active
huggingface/speech-to-speech
A low-latency, fully modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) built from open-source models, exposed via the OpenAI Realtime…
8512882active
PaddlePaddle/PaddleSpeech
PaddleSpeech is an open-source speech and audio toolkit built on the PaddlePaddle deep learning platform, covering ASR with punctuation, st…
6612670active
abus-aikorea/voice-pro
Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, …
8412647active
elebumm/RedditVideoMakerBot
A Python application that automatically generates Reddit story videos for platforms like TikTok, YouTube, and Instagram with a single comma…
7212532active
YiiGuxing/TranslationPlugin
A translation plugin for IntelliJ-based IDEs and Android Studio that supports multiple translation engines including Google, Microsoft, Dee…
9511831active
rany2/edge-tts
A Python module and CLI that lets you use Microsoft Edge's online neural text-to-speech service without needing Edge, Windows, or an API ke…
7911803active
debpalash/VoiceStudio
VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, vid…
8011737active
ankidroid/Anki-Android
AnkiDroid is the Android client for the open-source Anki spaced repetition flashcard system, letting users study and sync decks with AnkiWe…
9311628active
CorentinJ/Real-Time-Voice-Cloning
A Python implementation of the SV2TTS (Transfer Learning from Speaker Verification to Multispeaker TTS) framework that clones a voice from …
6460110maintenance
krillinai/KrillinAI
KrillinAI is an AI-powered video translation and dubbing tool covering the full pipeline: video download, speech transcription, subtitle tr…
8411272active
SparkAudio/Spark-TTS
Spark-TTS is an LLM-based text-to-speech system built on Qwen2.5 that generates speech directly from single-stream decoupled speech tokens.…
2711006active
kyutai-labs/moshi
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai, built around the Mimi streaming neural audio…
6210949active
moonshine-ai/moonshine
Moonshine Voice is an open-source on-device AI toolkit providing very low latency speech-to-text, intent recognition, and text-to-speech fo…
8910944active
linyqh/NarratoAI
NarratoAI is an open-source, self-hosted AI-powered video commentary and automated editing tool. It uses large language models to generate …
8810879active
NVIDIA/personaplex
PersonaPlex is a real-time, full-duplex speech-to-speech conversational model from NVIDIA that supports persona control via text role promp…
4710393active
niedev/RTranslator
RTranslator is a free, open-source, offline real-time translation app for Android that runs speech recognition (Whisper) and translation (M…
8510355active
RunanywhereAI/runanywhere-sdks
RunAnywhere is a set of cross-platform SDKs (Swift, Kotlin, React Native, Flutter, TypeScript, C++) over a shared C++ core for running AI m…
8510282active
open-mmlab/Amphion
Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s…
5110271active
espnet/espnet
ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance…
889941active
ripperhe/Bob
Bob is a macOS menu bar application for translation and OCR, supporting selection translation, screenshot translation, input translation, s…
499735active
k2-fsa/OmniVoice
OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style a…
799455active
mengxi-ream/read-frog
Read Frog is an open-source AI-powered browser extension for language learning and immersive translation. It translates web pages, selected…
829346active
coqui-ai/TTS
Coqui TTS is a deep learning toolkit for text-to-speech synthesis, providing pretrained models in over 1100 languages plus tools for traini…
2345953maintenance
kyutai-labs/pocket-tts
Pocket TTS is a lightweight 100M-parameter text-to-speech model and Python library from Kyutai that runs efficiently on CPUs without a GPU.…
839135active
jianchang512/clone-voice
A voice cloning tool with a web interface built on the coqui.ai xtts_v2 model, letting users synthesize speech in any voice from text or co…
108989active
jasonppy/VoiceCraft
VoiceCraft is a token infilling neural codec language model for zero-shot speech editing and text-to-speech on in-the-wild data like audiob…
638574active
hexgrad/kokoro
An inference library for Kokoro-82M, an open-weight text-to-speech model with 82 million parameters that delivers quality comparable to lar…
368570active
netease-youdao/EmotiVoice
EmotiVoice is an open-source text-to-speech engine supporting English and Chinese with over 2000 voices and prompt-controlled emotional syn…
178523active
santinic/audiblez
Audiblez is a Python CLI tool (with an optional GUI) that converts .epub e-books into .m4b audiobooks using the Kokoro-82M text-to-speech m…
528460active
duixcom/Duix-Mobile
Duix Mobile is an open-source SDK for building real-time interactive AI avatars (digital humans) that run on-device on Android, iOS, tablet…
698198active
GetStream/Vision-Agents
An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O…
848100active
suno-ai/bark
Bark is Suno's open-source transformer-based text-to-audio model that generates highly realistic multilingual speech, music, background noi…
3039249maintenance
RayVentura/ShortGPT
ShortGPT is a Python framework that automates AI-driven video content creation, combining LLM script generation, voiceover synthesis, asset…
337884active
STranslate/STranslate
STranslate is a ready-to-go Windows desktop translation and OCR tool built with WPF. It aggregates dozens of translation services (OpenAI, …
957860active
Blaizzy/mlx-audio
MLX-Audio is a Python library built on Apple's MLX framework for fast text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS…
887793active
jianchang512/ChatTTS-ui
A local web interface for the ChatTTS text-to-speech model that synthesizes speech from mixed Chinese/English text, numbers, and symbols. I…
707640active
babysor/MockingBird
MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real…
5436909maintenance
myshell-ai/MeloTTS
MeloTTS is a high-quality multi-lingual text-to-speech library supporting English (multiple accents), Spanish, French, Chinese, Japanese, a…
167608active
farzaa/clicky
Clicky is an open-source macOS app that acts as an AI teacher living next to your cursor — it can see your screen, talk with you via voice,…
507394active
Zyphra/Zonos
Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I…
247244active
wzpan/wukong-robot
wukong-robot is a modular Chinese-language voice assistant / smart speaker project in Python that combines offline wake-word detection, ASR…
327126active
ddean2009/MoneyPrinterPlus
A Python desktop application that uses AI LLMs to batch-generate short videos with one click, including automatic video mashup/remixing and…
307008active
yihong0618/xiaogpt
A Python application that connects Xiaomi AI Speakers to ChatGPT and other LLMs, letting users ask questions starting with a trigger phrase…
646901active
espeak-ng/espeak-ng
eSpeak NG is a compact open-source text-to-speech synthesizer supporting over 100 languages and accents, using formant synthesis for small …
676763active
microsoft/call-center-ai
An AI-powered call center service that lets you initiate or receive phone calls handled by an LLM-driven agent via a simple API call. Built…
646563active
souzatharsis/podcastfy
Podcastfy is an open-source Python package and CLI that transforms multimodal content (websites, PDFs, images, YouTube videos, topics) into…
606521active
argmaxinc/argmax-oss-swift
A Swift SDK providing turn-key on-device speech AI frameworks for Apple Silicon, including WhisperKit (speech-to-text with Whisper), Speake…
916338active
yl4579/StyleTTS2
StyleTTS 2 is a PyTorch text-to-speech model that uses style diffusion and adversarial training with large speech language models (e.g., Wa…
296336active
canopyai/Orpheus-TTS
Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a…
456314active
neuphonic/neutts
NeuTTS is a collection of open-source, on-device text-to-speech models built on small LLM backbones, with instant voice cloning from as lit…
606256active
Shaunwei/RealChar
RealChar is an open-source application for creating, customizing, and talking to AI characters/companions in realtime via voice or text. It…
486213active
bytedance/MegaTTS3
MegaTTS 3 is ByteDance's open-source PyTorch text-to-speech model with a lightweight 0.45B-parameter Diffusion Transformer backbone. It pro…
596091active
HapeLee/legado-with-MD3
Legado with MD3 is a free, open-source Android e-book and content aggregation reader, rebuilt from the Legado (阅读 3.0) project with a Mater…
825860active
denizsafak/abogen
Abogen is a desktop text-to-speech application that converts EPUB, PDF, text, markdown, and subtitle files into audiobooks with synchronize…
755771active
dnhkng/GLaDOS
A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,…
645689active
huggingface/parler-tts
Parler-TTS is a lightweight text-to-speech library from Hugging Face that generates high-quality, natural-sounding speech controllable via …
255586active
dograh-hq/dograh
Dograh is an open-source, self-hostable voice AI platform for building production voice agents, positioned as an alternative to Vapi and Re…
845505active
remsky/Kokoro-FastAPI
A Dockerized FastAPI wrapper around the Kokoro-82M text-to-speech model exposing an OpenAI-compatible speech endpoint with CPU, NVIDIA, AMD…
885373active
liuzhao1225/YouDub-webui
YouDub WebUI is an open-source AI video localization and dubbing tool that converts YouTube, Bilibili, or local videos into target-language…
755358active
ysharma3501/LuxTTS
LuxTTS is a lightweight zipvoice-based text-to-speech model for high-quality zero-shot voice cloning, generating clear 48kHz speech at up t…
545306active
OHF-Voice/piper1-gpl
Piper is a fast, fully local neural text-to-speech engine that embeds espeak-ng for phonemization and ships with a CLI, HTTP web server, Py…
865287active
YILS-LIN/short-video-factory
Short Video Factory is an open-source cross-platform desktop application that uses AI to generate and automatically edit short videos from …
785176active
Plachtaa/VITS-fast-fine-tuning
A Python pipeline for fast fine-tuning of VITS text-to-speech models, enabling speaker adaptation in under an hour from short audio, long a…
105012active
buxuku/SmartSub
SmartSub (妙幕) is a free, open-source cross-platform desktop application that provides an end-to-end subtitle and dubbing pipeline: speech-t…
894778active
WhisperSpeech/WhisperSpeech
WhisperSpeech is an open-source text-to-speech system built by inverting OpenAI's Whisper model, aiming to be 'Stable Diffusion for speech'…
564639active
gradio-app/fastrtc
FastRTC is a Python library that turns any Python function into a real-time audio and video stream over WebRTC or WebSockets. It includes b…
584622active
jing332/tts-server-android
An Android system-wide text-to-speech (TTS) app that provides a TTS engine service with built-in Microsoft demo API support, custom HTTP TT…
404483active
pot-app/pot-desktop
Pot is a cross-platform desktop application for hotkey-based text translation, screenshot OCR, and text-to-speech, built with Tauri. It sup…
6219343maintenance
dramaclaw/dramaclaw
DramaClaw is a self-hosted, general-purpose AIGC video engine that turns scripts into finished films through a single pipeline covering sto…
814355active

page 1 / 5 next →