Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: tts

498 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
MOSS-TTS
MOSS-TTS-Nano is an open-source 0.1B-parameter multilingual speech generation (TTS) model from MOSI.AI and the OpenMOSS team, designed for …
574031active
KoljaB/RealtimeTTS
RealtimeTTS is a Python library that converts text, generators, and LLM token streams into speech audio with low latency. It supports multi…
944016active
QwenLM/Qwen3-Omni
Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an…
523980active
mushan0x0/AI0x0.com
AI 0x0 is a multimodal, multi-model desktop AI assistant that lives as a floating ball on macOS and Windows, letting users invoke AI querie…
343938active
omnivore-app/omnivore
Omnivore is a complete, open source read-it-later application for saving and reading articles, with highlighting, notes, search, labels, PD…
8816226maintenance
PeterH0323/Streamer-Sales
Streamer-Sales is a fine-tuned LLM-based AI sales livestreamer application that generates persuasive product commentary from product descri…
273761active
HumanAIGC-Engineering/OpenAvatarChat
OpenAvatarChat is a modular interactive digital human (talking avatar) chat application that combines ASR, LLM, TTS, and avatar rendering c…
743723active
speaches-ai/speaches
Speaches is an OpenAI API-compatible self-hosted server for speech-to-text (via faster-whisper), translation, and text-to-speech (via Kokor…
783621active
Soul-AILab/SoulX-Podcast
SoulX-Podcast is the official inference codebase for a text-to-speech model that generates long-form, multi-turn, multi-speaker podcast-sty…
423535active
neonbjb/tortoise-tts
Tortoise TTS is a multi-voice text-to-speech library built on PyTorch that prioritizes highly realistic prosody and intonation. It combines…
3214870maintenance
Kedreamix/Linly-Talker
Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, …
483436active
Kedreamix/Linly-Dubbing
Linly-Dubbing is an intelligent multi-language AI dubbing and video translation tool that combines speech recognition (WhisperX, FunASR), L…
273331active
rsxdalv/TTS-WebUI
TTS WebUI is a free, open-source web interface combining Gradio and React that unifies 30+ AI models for text-to-speech, voice conversion, …
893244active
VOICEVOX
VOICEVOX is a free, mid-quality text-to-speech (and singing synthesis) software whose editor is built with Electron, TypeScript, and Vue. I…
933231active
intel/acat
ACAT is an open-source assistive communication platform from Intel Labs, originally developed for Stephen Hawking, that helps people with r…
653230active
MisoLabsAI/MisoTTS
Miso TTS 8B is an open-source text-to-speech model based on an RVQ Transformer architecture with a Llama 3.2-style 8B backbone, designed fo…
523224active
zai-org/GLM-4-Voice
GLM-4-Voice is an end-to-end bilingual (Chinese/English) speech dialogue model from Zhipu AI, built on GLM-4-9B with a speech tokenizer and…
223222active
BayLing-Models/BayLing-Speech
LLaMA-Omni is an end-to-end speech interaction model built on Llama-3.1-8B-Instruct that generates simultaneous text and speech responses f…
323146active
SuperCmdLabs/SuperCmd
SuperCmd is an open-source macOS launcher combining Raycast-compatible extensions, hold-to-speak dictation, text-to-speech, AI chat and age…
753124active
HeyWillow/willow
Willow is an open source, self-hosted voice assistant platform for ESP32-S3-BOX hardware, designed as a privacy-focused alternative to Amaz…
863097active
theajack/cnchar
cnchar is a comprehensive TypeScript library for Chinese character processing, offering pinyin conversion, stroke counts, stroke order draw…
663083active
elevenlabs/elevenlabs-python
The official Python SDK for the ElevenLabs API, providing programmatic access to text-to-speech, speech-to-text, voice cloning, dubbing, mu…
933078active
kyutai-labs/delayed-streams-modeling
Kyutai's repository of Speech-To-Text and Text-To-Speech models built on the Delayed Streams Modeling framework, with implementations in Py…
473017active
off-grid-ai/OGAM
Off Grid AI (OGAM) is a cross-platform mobile and desktop application that runs AI entirely on-device: GGUF LLM chat with vision, Whisper s…
783002active
KevinWang676/Bark-Voice-Cloning
A one-click hub of Gradio Web UIs and Colab notebooks for open-source voice cloning, TTS, and voice conversion models including Bark, GPT-S…
672947active
openai/openai-fm
OpenAI.fm is an interactive web demo showcasing OpenAI's text-to-speech models, built with Next.js and the OpenAI Speech API. It lets users…
512887active
Camb-ai/MARS5-TTS
MARS5 is an open-source English text-to-speech model from CAMB.AI that uses a two-stage AR-NAR pipeline to generate expressive speech with …
142817active
lnreader/lnreader
LNReader is a free, open-source light novel and webnovel reader app for Android, inspired by Tachiyomi. It supports 200+ community-maintain…
872793active
rhasspy/piper
Piper is a fast, local neural text-to-speech system that runs offline on modest hardware, including Raspberry Pi devices. It offers many pr…
1011281maintenance
AutoArk/GPA
GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an…
542762active
hahahumble/speechgpt
SpeechGPT is an open-source web application that lets users have voice conversations with ChatGPT using speech recognition and speech synth…
622751active
quik-sms/quik
QUIK is an open-source SMS messenger app for Android, a revived continuation of QKSMS. It replaces the stock messaging app with features li…
872708active
FluidInference/FluidAudio
A Swift SDK providing fully local, low-latency audio AI on Apple devices, including speech-to-text (Parakeet), speaker diarization (Pyannot…
812699active
thewh1teagle/kokoro-onnx
A Python library that runs the Kokoro text-to-speech model via ONNX Runtime, supporting CPU and GPU inference. It provides multi-language T…
672680active
ZeframLou/call-me
A minimal Claude Code plugin that places real phone calls to notify you when an AI coding agent finishes a task, gets stuck, or needs a dec…
542638active
pndurette/gTTS
gTTS is a Python library and CLI tool that interfaces with Google Translate's text-to-speech API to generate spoken MP3 audio from text. It…
572627stable
nvaccess/nvda
NVDA (NonVisual Desktop Access) is a free, open source screen reader for Microsoft Windows that reads on-screen text aloud via synthetic sp…
902625stable
6drf21e/ChatTTS_colab
A one-click deployment wrapper around ChatTTS providing a Gradio web UI for text-to-speech, runnable in Google Colab or via an offline Wind…
532592active
liou666/polyglot
Polyglot is a cross-platform desktop (and web) application for practicing spoken language with AI conversation partners, built on ChatGPT f…
392585active
ahmedeltaher/Android-MVVM-Architecture-Android-Voice-AI-SDK
A reusable Android library (Kotlin, MVVM) that provides a full voice-driven AI conversation pipeline: microphone capture with VAD, speech-t…
702578active
miantiao-me/hacker-podcast
An AI-powered Chinese podcast application that automatically scrapes daily Hacker News top stories, generates Chinese summaries and scripts…
622576active
zixiiu/Digital_Life_Server
A Python server backend for a 'digital life' voice assistant that combines speech recognition, ChatGPT-based conversation, sentiment analys…
302554active
AIGC-Audio/AudioGPT
AudioGPT is a Python framework that wraps multiple audio foundation models (for speech, singing, sound, and talking-head tasks) behind a GP…
3010167maintenance
mozilla/TTS
A deep learning library for advanced text-to-speech generation, built on PyTorch with models like Tacotron2, Glow-TTS, and various vocoders…
2310167maintenance
nateshmbhat/pyttsx3
pyttsx3 is an offline text-to-speech synthesis library for Python 3 that wraps system TTS engines like Sapi5, NSSpeechSynthesizer, and espe…
672529active
janhq/ichigo
Ichigo is a Python speech package for developers offering local realtime voice AI capabilities, including a compact 22M-parameter speech to…
502492active
erew123/alltalk_tts
AllTalk TTS is a text-to-speech application built on the Coqui TTS engine, usable standalone or as an extension for Text-generation-webui, …
442429active
pnnbao97/VieNeu-TTS
VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi…
842427active
wan-h/awesome-digital-human-live2d
An open-source digital human application that combines Live2D avatars with LLM-powered conversation, integrating ASR, LLM, TTS, and agent o…
622413active
showlab/Paper2Video
Paper2Video is a Python pipeline that automatically generates academic presentation videos from scientific papers, taking a paper PDF, a sp…
482368active
jik876/hifi-gan
The official PyTorch implementation of HiFi-GAN, a generative adversarial network that converts mel-spectrograms into high-fidelity 22.05 k…
322367stable
Mentra-Community/MentraOS
MentraOS is an open-source operating system and development platform for smart glasses, providing pairing, connection management, data stre…
942326active
cosin2077/easyVoice
EasyVoice is an open-source text-to-speech application that converts long texts and novels into high-quality audio with streaming playback …
492287active
QwenAudio/qwen-audio-agent
A realtime voice runtime and frontend for AI coding agents like Claude Code, Codex, and Qwen Code, letting agents talk, listen, and report …
792264active
fishaudio/Bert-VITS2
Bert-VITS2 is a text-to-speech model implementation combining the VITS2 architecture with multilingual BERT embeddings, written in Python. …
638796maintenance
rushindrasinha/youtube-shorts-pipeline
Verticals v3 (repo youtube-shorts-pipeline) is a Python CLI that automates producing and publishing YouTube Shorts: it researches a topic, …
612251active
lifeiteng/vall-e
An unofficial PyTorch implementation of VALL-E, a zero-shot text-to-speech model that treats TTS as a conditional language modeling task ov…
402215active
DigitalPhonetics/IMS-Toucan
IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive…
632207active
boson-ai/higgs-audio
Higgs Audio is a text-audio foundation model project from Boson AI providing code and weights for conversational text-to-speech with zero-s…
578329maintenance
MiniMax-AI/cli
The official CLI for the MiniMax AI Platform, written in TypeScript, that generates text, images, video, speech, and music from the termina…
812073active
kimjammer/Neuro
A local recreation of the Neuro-Sama AI VTuber that runs open-source LLMs on consumer hardware, combining realtime speech-to-text, text-to-…
282070active
travisvn/openai-edge-tts
A self-hosted Python service that emulates the OpenAI text-to-speech API endpoint (/v1/audio/speech) using Microsoft Edge's free online TTS…
262066active
Plachtaa/VALL-E-X
An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint…
107931maintenance
jaywalnut310/vits
VITS is the official PyTorch implementation of an end-to-end text-to-speech model based on a conditional variational autoencoder with adver…
327889maintenance
0xShug0/audio.cpp
audio.cpp is a pure C++ inference engine for audio models built on ggml, supporting TTS, STT, VAD, voice conversion, music generation, and …
802022active
digitalsamba/claude-code-video-toolkit
An AI-native video production toolkit designed for Claude Code, providing skills, commands, templates, and Python tools so an AI agent can …
832008active
run-llama/notebookllama
NotebookLlaMa is an open-source, Python-based alternative to Google's NotebookLM, backed by LlamaCloud for document ingestion and retrieval…
521967active
praat/praat.github.io
Praat is a desktop application for analyzing, synthesizing, and manipulating speech, widely used in phonetics research and teaching. It pro…
991965stable
akdeb/ElatoAI
ElatoAI is an open-source platform for running realtime voice AI conversations on ESP32/Arduino hardware, supporting 100+ STT, LLM, and TTS…
641933active
AlexxIT/YandexStation
A Home Assistant custom component (installed via HACS) for controlling Yandex smart speakers and other Alice smart home devices. It support…
971917active
flybirdxx/ComfyUI-Qwen-TTS
A ComfyUI custom node plugin that wraps Alibaba's Qwen3-TTS model for speech synthesis, zero-shot voice cloning, and natural-language voice…
541876active
MixLabPro/comfyui-mixlab-nodes
A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec…
671863active
geometer/FBReaderJ
FBReaderJ is the official repository behind FBReader, a long-running e-book reader application for Android written in Java. It supports for…
321849active
RHVoice/RHVoice
RHVoice is a free and open-source statistical parametric speech synthesizer (TTS) built on HTS technology, originally for Russian and now s…
911831active
nazdridoy/kokoro-tts
A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w…
881816active
k2-fsa/sherpa-ncnn
A C++ library for real-time offline speech recognition, text-to-speech, and voice activity detection built on the ncnn inference framework …
581779active
wwbin2017/bailing
Bailing is an open-source voice assistant application similar to GPT-4o, built with an ASR + VAD + LLM + TTS pipeline (FunASR, silero-vad, …
541757active
High-Logic/Genie-TTS
GENIE is a lightweight Python inference engine for the open-source GPT-SoVITS text-to-speech project, optimized for fast CPU-based speech s…
611754active
ThioJoe/Auto-Synced-Translated-Dubs
A Python CLI tool that automatically translates video subtitles into multiple languages and generates AI voice dubbed audio tracks synced t…
591747active
ken107/read-aloud
Read Aloud is a browser extension for Chrome, Firefox, and Edge that reads aloud webpage content using text-to-speech with one click. It wo…
641731active
trymirai/uzu
Uzu is a high-performance inference engine written in Rust for running AI models directly on-device, with Python, TypeScript, and Swift bin…
851678active
LokerL/tts-vue
TTS-Vue is a cross-platform desktop text-to-speech application built with Electron, Vue, ElementPlus, and Vite that uses Microsoft's Edge T…
586101maintenance
mkiol/dsnote
Speech Note is a Linux desktop and Sailfish OS application for note taking, reading, and translating text using offline Speech to Text, Tex…
961607active
Norsico/Video-Materials-AutoGEN-Workstation
A self-hosted short-video production workstation that combines AI script generation (Gemini), batch TTS voiceover, AI image asset synthesis…
541599active
semperai/amica
Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. …
321593active
Alisa0808/vox-director
An agent skill that turns a single topic into a finished Vox-style paper-collage explainer or ad video, automating script, keyframes, motio…
561582active
Agents365-ai/video-podcast-maker
An agent skill/plugin that turns a plain-language topic into a 4K narrated video podcast, combining research, script generation, multi-engi…
791580active
MiniMax-AI/MiniMax-MCP
The official MiniMax Model Context Protocol (MCP) server, written in Python, exposing MiniMax's text-to-speech, image generation, and video…
641569active
hanmin0822/MisakaTranslator
MisakaTranslator is a Windows desktop application that provides real-time machine translation for Galgames, text-based games, and manga. It…
255746maintenance
Enemyx-net/VibeVoice-ComfyUI
A ComfyUI custom node integration for Microsoft's VibeVoice text-to-speech model, providing single and multi-speaker voice synthesis with v…
531549active
RunanywhereAI/RCLI
RCLI is a single-binary CLI that runs open-source AI models locally on your machine, covering chat, vision, speech-to-text, text-to-speech,…
821542active
vibevoice-community/VibeVoice
VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker …
601542active
dindin0497/SeeIt
SeeIt is an inclusive Android app with two accessibility modes: one that converts spoken speech into text and plays corresponding ASL (Amer…
411540active
AlekPet/ComfyUI_Custom_Nodes_AlekPet
A collection of custom nodes for ComfyUI that extend its capabilities with painting, pose control, prompt translation, and speech recogniti…
721524active
kyutai-labs/unmute
Unmute is a system that lets any text LLM listen and speak by wrapping it with Kyutai's low-latency speech-to-text and text-to-speech model…
601506active
met4citizen/TalkingHead
A JavaScript library for real-time lip-synced talking avatars, rendering full-body 3D characters in the browser. It supports text-to-speech…
671505active
stepfun-ai/Step-Audio2
Step-Audio 2 is an end-to-end multimodal large language model for industry-strength audio understanding and speech conversation, with open-…
501503active
espressif/esp-sr
ESP-SR is Espressif's speech recognition framework for ESP32-series chips, providing wake word detection (WakeNet), voice activity detectio…
781492active
ekwek1/soprano
Soprano is an ultra-lightweight 80M-parameter text-to-speech model and Python library for fast, expressive, high-fidelity speech synthesis …
441486active
NVIDIA/tacotron2
NVIDIA's PyTorch implementation of the Tacotron 2 text-to-speech model, which synthesizes mel spectrograms from text for vocoder-based audi…
325296maintenance

← prev page 2 / 5 next →