Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: tts

498 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
microsoft/NeuralSpeech
NeuralSpeech is a Microsoft Research Asia research repository containing implementations of neural speech processing models across ASR erro…
321462active
869413421/ai-moive-studio
AICON is a full-stack, self-hostable AI video creation studio that pairs a natural-language-driven agent with an infinite-canvas node edito…
601456active
DicioTeam/dicio-android
Dicio is a free and open source voice assistant app for Android that interprets user questions and provides speech and graphical feedback e…
831452active
sauravpanda/BrowserAI
BrowserAI is a TypeScript library for running LLMs, speech recognition, text-to-speech, and audio separation models directly in the browser…
781449active
voice-cloning-app/Voice-Cloning-App
A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom…
231440active
edwko/OuteTTS
OuteTTS is a Python interface for running OuteAI's text-to-speech models, supporting multiple inference backends including llama.cpp, Huggi…
521436active
FireRedTeam/FireRedTTS2
FireRedTTS-2 is a long-form streaming text-to-speech system for multi-speaker dialogue generation, built in PyTorch with a dual-transformer…
401428active
devnen/Chatterbox-TTS-Server
A self-hosted server wrapping Resemble AI's Chatterbox TTS models behind an OpenAI-compatible API with a modern web UI. It supports voice c…
651420active
Voine/ChatWaifu_Mobile
An Android app that lets users chat with an anime-style AI companion powered by ChatGPT, with local VITS text-to-speech, Live2D character r…
641419active
lenML/Speech-AI-Forge
Speech-AI-Forge is a Python project built around multiple TTS generation models (ChatTTS, CosyVoice, Fish-Speech, Index-TTS, F5-TTS, FireRe…
621416active
R3gm/SoniTranslate
SoniTranslate is a Gradio-based web application that automatically dubs videos into other languages. It transcribes speech, translates it, …
551410active
0nutation/SpeechGPT
SpeechGPT is a series of speech large language models that can perceive and generate speech, supporting cross-modal instruction following a…
291401active
lamm-mit/PDF2Audio
A Gradio-based web application that converts PDF documents into audio podcasts, lectures, and summaries using OpenAI GPT models for text ge…
301383active
ssnangua/ColorTxt
ColorTxt is a local desktop novel reader for TXT and common ebook formats (epub, mobi, pdf, etc.) that colorizes content with custom highli…
801374active
elevenyellow/handcrafted-persona-engine
Persona Engine is a Windows desktop application that drives a Live2D avatar with an AI pipeline: microphone speech recognition, an LLM guid…
731357active
shivammehta25/Matcha-TTS
Matcha-TTS is a PyTorch-based text-to-speech system that uses conditional flow matching for fast, non-autoregressive speech synthesis. It s…
621349active
mbailey/voicemode
VoiceMode is a Python-based MCP server and Claude Code plugin that enables natural, hands-free voice conversations with Claude Code and oth…
801339active
google-gemma/gemma-translator
An on-device, fully offline voice translator application powered by Gemma 4 via LiteRT-LM, with a React web frontend optimized for small ha…
561337active
foobnix/LibreraReader
Librera Reader is a highly customizable e-book reader application for Android supporting PDF, EPUB, MOBI, DjVu, FB2, CBZ/CBR and many other…
974740maintenance
andimarafioti/faster-qwen3-tts
A Python library for real-time text-to-speech inference with Qwen3-TTS using manual CUDA graph capture, requiring no Flash Attention, vLLM,…
781330active
Capsize-Games/airunner
AI Runner is a privacy-focused desktop application for running local AI models offline, combining an AI chat companion with voice conversat…
791314active
gyoridavid/short-video-maker
An open-source server that generates short-form videos (TikTok, Instagram Reels, YouTube Shorts) from text prompts, combining Kokoro TTS, a…
321311active
Henry-23/VideoChat
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late…
481303active
ttop32/MouseTooltipTranslator
A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports …
951300active
panyanyany/Twocast
Twocast is a self-hostable AI podcast generator that creates two-person conversational podcast episodes from topics, links, documents, or w…
321290active
ictnlp/StreamSpeech
StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p…
371287active
jasperproject/jasper-client
Client code for the Jasper voice computing platform, an open source platform for building always-on, voice-controlled applications. It prov…
324518maintenance
perminder-klair/subwave
SUB/WAVE is a self-hostable personal internet radio station that broadcasts a single shared Icecast stream to all listeners simultaneously.…
771278active
studio-dots-ai/dots.tts
dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive text-to-speech system, distributed as a Python library with pretrain…
791275active
phuc-nt/my-translator
A Tauri-based desktop app that captures system or microphone audio, transcribes it, and shows translations in a minimal real-time overlay, …
761275active
ABexit/ASR-LLM-TTS
An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) …
601271active
gitmylo/audio-webui
An all-in-one web UI for audio-related neural networks, bundling text-to-speech (Bark), voice conversion/cloning (RVC), and text-to-audio/m…
301247active
sh-lee-prml/HierSpeechpp
Official PyTorch implementation of HierSpeech++, a fast zero-shot speech synthesizer for text-to-speech and voice conversion based on hiera…
281238active
Ksuriuri/index-tts-vllm
A reimplementation of IndexTTS's GPT model inference using vLLM, providing significantly faster text-to-speech generation with a web UI and…
581235active
Aratako/Irodori-TTS
Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer…
581219active
uezo/ChatdollKit
ChatdollKit is a Unity SDK that turns 3D character models into voice-enabled chatbots and virtual assistants. It integrates LLMs (ChatGPT, …
761213active
hgneng/ekho
Ekho is an open-source Chinese text-to-speech engine supporting Mandarin, Cantonese, and Tibetan, part of the eGuideDog accessibility proje…
681211active
metavoiceio/metavoice-src
MetaVoice-1B is a 1.2B parameter foundational text-to-speech model trained on 100K hours of speech, focused on emotional rhythm and tone in…
264205maintenance
hkjarral/AVA-AI-Voice-Agent-for-Asterisk
An open-source AI voice agent that integrates with Asterisk/FreePBX phone systems via Audiosocket/RTP, built in Python with a modular pipel…
851202active
kizuna-ai-lab/sokuji
Sokuji is a cross-platform real-time two-way speech translation app for bilingual meetings, available as a desktop application (Windows, ma…
831199active
diodiogod/TTS-Audio-Suite
A ComfyUI custom node suite providing unified multi-engine Text-to-Speech, Voice Conversion, and audio editing across 19 engines like Chatt…
841185active
nari-labs/dia2
Dia2 is a streaming dialogue text-to-speech model from Nari Labs that can generate conversational audio in realtime without needing the ful…
411172active
DavidVentura/offline-translator
An Android app that translates text, PDF/ODT documents, and images entirely offline using Firefox translation models on-device. It also off…
881166active
soniqo/speech-swift
An open-source Swift toolkit for on-device speech AI on Apple Silicon, providing ASR, TTS, speech-to-speech, VAD, and speaker diarization v…
831161active
alibaba/lumenx
LumenX is an AI-native platform for turning novel text into publishable motion comic and short drama videos. It provides a full pipeline fr…
591158active
janvarev/Irene-Voice-Assistant
Irene is an offline-capable Russian voice assistant written in Python that recognizes speech via Vosk and responds using TTS engines. Its f…
651152active
TensorSpeech/TensorFlowTTS
TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me…
233995maintenance
PromtEngineer/Verbi
Verbi is a modular Python voice assistant application for experimenting with state-of-the-art transcription, LLM response generation, and t…
481124active
steveseguin/social_stream
Social Stream Ninja is a free, open-source tool that consolidates live chat messages from 120+ social platforms (YouTube, Twitch, TikTok, F…
1001121active
FutureUniant/Tailor
Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea…
371115active
ardha27/AI-Waifu-Vtuber
An AI waifu Vtuber application that listens to your voice via Whisper speech recognition, generates in-character replies with OpenAI, and s…
681113active
KoljaB/RealtimeVoiceChat
A Python client-server application enabling natural spoken conversations with LLMs, streaming browser audio via WebSockets through Realtime…
333832maintenance
LoredCast/filewizard
A self-hosted, browser-based web UI for converting files between many formats, running OCR on PDFs and images, transcribing audio with Whis…
461102active
vocodedev/vocode-core
Vocode is an open-source Python library for building real-time, voice-based LLM applications and agents. It orchestrates streaming transcri…
213785maintenance
nobodywho-ooo/nobodywho
NobodyWho is an open-source (EUPL-1.2) on-device LLM inference engine written in Rust, built on llama.cpp, with SDKs for Kotlin, Swift, Pyt…
861086active
WhiskeyCoder/Qwen3-Audiobook-Converter
A Python CLI tool that converts documents (PDF, EPUB, DOCX, DOC, TXT) into audiobooks using the Qwen3 TTS voice model running locally via a…
491081active
festvox/flite
Flite (festival-lite) is a small, fast, portable run-time text-to-speech synthesis engine written entirely in ANSI C by Carnegie Mellon Uni…
231081stable
IAHispano/Applio
Applio is an open-source, MIT-licensed voice conversion suite built on RVC that lets users convert audio into other voices, train custom vo…
903644maintenance
brenpoly/be-more-agent
An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp…
641060active
Makememo/MemoAI
MemoAI is a desktop application for macOS and Windows that transcribes audio and video (YouTube links, podcasts, local files) into text and…
961059active
C-Loftus/QuickPiperAudiobook
A Go CLI tool that converts text content from formats like epub, PDF, mobi, txt, HTML, and docx into natural-sounding audiobooks with a sin…
491057active
zai-org/GLM-TTS
GLM-TTS is a Python-based text-to-speech synthesis system built on large language models, using a two-stage LLM plus Flow model architectur…
501055active
tegnike/aituber-kit
AITuberKit is an all-in-one web application toolkit for building and deploying AI character chat experiences, including streaming-oriented …
891054active
translate-tools/linguist
Linguist is a privacy-first browser extension for Chrome and Firefox that translates web pages, selected text, subtitles, and messages, wit…
981053active
wxxxcxx/ms-ra-forwarder
A self-hostable free online text-to-speech API that proxies Microsoft Edge 'Read Aloud' and Azure TTS demo endpoints. It can be deployed vi…
611053active
Artrajz/vits-simple-api
A Python HTTP API service that exposes VITS-family text-to-speech models (VITS, Bert-VITS2, GPT-SoVITS, W2V2 emotional VITS) for inference.…
691052active
shang-zhu/violin
Violin is an open-source video translation tool that transcribes speech, translates it into 33 languages, synthesizes a native-sounding voi…
521047active
k2-fsa/ZipVoice
ZipVoice is a series of fast, high-quality zero-shot text-to-speech models based on flow matching, with a compact 123M-parameter Zipformer-…
431045active
xiaochong/hi-kid
HiKid is a free, open-source Electron desktop app that lets children in non-English-speaking countries practice English speaking and listen…
541043active
PatterAI/Patter
Patter is an open-source, MIT-licensed SDK (Python and TypeScript) that connects AI agents to real phone calls, handling telephony, speech-…
781039active
Goekdeniz-Guelmez/Local-NotebookLM
A local, open-source alternative to Google's NotebookLM that converts PDF documents into audio content like podcasts, summaries, and interv…
741027active
mybigday/llama.rn
A React Native binding of llama.cpp that enables on-device LLM inference on iOS and Android. It supports GPU/NPU acceleration (Metal, Hexag…
941026active
rayenfeng/riko_project
Project Riko is an anime-themed conversational voice assistant that combines OpenAI's GPT for dialogue, GPT-SoVITS for voice synthesis, and…
311023active
jianjieyiban/JJYB_AI_VideoAutoCut
JJYB_AI 智剪 is a local-first desktop AI video creation workbench that combines material analysis, smart shot segmentation, commentary script…
701013active
HumeAI/tada
TADA is an open-source speech-language model from Hume AI that generates expressive, high-fidelity speech via text-acoustic dual alignment,…
511009active
InsiderX-Pro/video-translator
An OpenClaw agent skill (written in Shell/Python) that translates and dubs videos by submitting jobs to a remote video-translation service …
541007active
AutoArk/open-audio-opd
An industrial training stack for online policy distillation (OPD) of audio models, distilling compact ASR (and planned TTS) student models …
521007active
Jingyi-Wu-Richael/rachel-digital-human-production
A Codex skill (plugin) that packages a repeatable workflow for producing authorized digital-human talking-head videos using MiniMax voice c…
541005active
linyiLYi/bilibot
A local chatbot fine-tuned from Bilibili user comments, built on Qwen1.5-32B-Chat using Apple's MLX LoRA fine-tuning. It supports text chat…
243139maintenance
keithito/tacotron
An unofficial open-source TensorFlow implementation of Google's Tacotron end-to-end neural text-to-speech model, with a pre-trained model a…
232995maintenance
readbeyond/aeneas
aeneas is a Python/C library and set of CLI tools that automatically computes forced alignments, generating a synchronization map between a…
652863maintenance
FolioReader/FolioReaderKit
FolioReaderKit is a Swift framework for iOS that reads and parses ePub 2 and ePub 3 files, providing a full in-app ebook reader experience.…
102683maintenance
marytts/marytts
MaryTTS is an open-source, multilingual text-to-speech synthesis platform written in pure Java, operating as a client-server system. It sup…
242583maintenance
CjangCjengh/MoeGoe
MoeGoe is an executable command-line tool for running inference with VITS text-to-speech models, supporting TTS, voice conversion, HuBERT-V…
232423maintenance
cogentapps/chat-with-gpt
An open-source, self-hostable ChatGPT web app built with TypeScript and React that adds voice features via ElevenLabs text-to-speech and Op…
302352maintenance
NVIDIA/waveglow
WaveGlow is a PyTorch implementation of a flow-based generative network that synthesizes high-quality speech audio from mel-spectrograms, c…
322339maintenance
Rayhane-mamah/Tacotron-2
A TensorFlow implementation of DeepMind's Tacotron-2 neural text-to-speech architecture, including both the Tacotron spectrogram predictor …
322322maintenance
fatchord/WaveRNN
A PyTorch implementation of DeepMind's WaveRNN neural vocoder plus a Tacotron text-to-speech system, trained on LJSpeech. It supports train…
322190maintenance
ming024/FastSpeech2
A PyTorch implementation of Microsoft's FastSpeech 2 text-to-speech model, supporting English and Mandarin with single- and multi-speaker s…
322185maintenance
r9y9/deepvoice3_pytorch
A PyTorch implementation of Deep Voice 3 and related convolutional neural network-based text-to-speech synthesis models. It includes traini…
231975maintenance
PandaOCR
PandaOCR is a free Windows desktop OCR tool that captures screen regions and recognizes text using many cloud OCR engines (Sogou, Tencent, …
801918maintenance
Kyubyong/tacotron
A heavily documented TensorFlow implementation of Tacotron, a fully end-to-end text-to-speech synthesis model. It includes training, prepro…
321832maintenance
kalliope-project/kalliope
Kalliope is a modular, always-on voice-controlled personal assistant framework written in Python, designed to run on Linux systems includin…
231772maintenance
yakGPT/yakGPT
YakGPT is a locally running, browser-based ChatGPT UI that connects directly to the OpenAI API with your own key. It adds hands-free voice …
301586maintenance
Marak/say.js
A Node.js library that provides text-to-speech by shelling out to platform-native TTS engines (macOS `say`, Windows SAPI, Linux Festival). …
321531maintenance
s-macke/SAM
SAM (Software Automatic Mouth) is a tiny text-to-speech synthesizer written in C, adapted from the 1982 Commodore 64 speech software. It in…
321497maintenance
fossasia/MMM-SUSI-AI
A MagicMirror² module that integrates the SUSI.AI assistant, providing voice-activated intelligent answers on a smart mirror. It supports h…
321482maintenance
microsoft/SpeechT5
Microsoft's unified-modal speech-text pre-training framework implementing SpeechT5 and related models (SpeechLM, SpeechUT, VALL-E X, WavLLM…
321449maintenance
DragonComputer/Dragonfire
Dragonfire is an open-source virtual assistant for Ubuntu-based Linux distributions, combining speech recognition, text-to-speech, and NLP …
231407maintenance
innnky/emotional-vits
Emotional VITS is a voice synthesis (TTS) model based on VITS that enables emotion-controllable speech generation without requiring manual …
321392maintenance

← prev page 3 / 5 next →