function: tts
498 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| MOSS-TTS MOSS-TTS-Nano is an open-source 0.1B-parameter multilingual speech generation (TTS) model from MOSI.AI and the OpenMOSS team, designed for … | 57 | 4031 | active |
| KoljaB/RealtimeTTS RealtimeTTS is a Python library that converts text, generators, and LLM token streams into speech audio with low latency. It supports multi… | 94 | 4016 | active |
| QwenLM/Qwen3-Omni Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an… | 52 | 3980 | active |
| mushan0x0/AI0x0.com AI 0x0 is a multimodal, multi-model desktop AI assistant that lives as a floating ball on macOS and Windows, letting users invoke AI querie… | 34 | 3938 | active |
| omnivore-app/omnivore Omnivore is a complete, open source read-it-later application for saving and reading articles, with highlighting, notes, search, labels, PD… | 88 | 16226 | maintenance |
| PeterH0323/Streamer-Sales Streamer-Sales is a fine-tuned LLM-based AI sales livestreamer application that generates persuasive product commentary from product descri… | 27 | 3761 | active |
| HumanAIGC-Engineering/OpenAvatarChat OpenAvatarChat is a modular interactive digital human (talking avatar) chat application that combines ASR, LLM, TTS, and avatar rendering c… | 74 | 3723 | active |
| speaches-ai/speaches Speaches is an OpenAI API-compatible self-hosted server for speech-to-text (via faster-whisper), translation, and text-to-speech (via Kokor… | 78 | 3621 | active |
| Soul-AILab/SoulX-Podcast SoulX-Podcast is the official inference codebase for a text-to-speech model that generates long-form, multi-turn, multi-speaker podcast-sty… | 42 | 3535 | active |
| neonbjb/tortoise-tts Tortoise TTS is a multi-voice text-to-speech library built on PyTorch that prioritizes highly realistic prosody and intonation. It combines… | 32 | 14870 | maintenance |
| Kedreamix/Linly-Talker Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, … | 48 | 3436 | active |
| Kedreamix/Linly-Dubbing Linly-Dubbing is an intelligent multi-language AI dubbing and video translation tool that combines speech recognition (WhisperX, FunASR), L… | 27 | 3331 | active |
| rsxdalv/TTS-WebUI TTS WebUI is a free, open-source web interface combining Gradio and React that unifies 30+ AI models for text-to-speech, voice conversion, … | 89 | 3244 | active |
| VOICEVOX VOICEVOX is a free, mid-quality text-to-speech (and singing synthesis) software whose editor is built with Electron, TypeScript, and Vue. I… | 93 | 3231 | active |
| intel/acat ACAT is an open-source assistive communication platform from Intel Labs, originally developed for Stephen Hawking, that helps people with r… | 65 | 3230 | active |
| MisoLabsAI/MisoTTS Miso TTS 8B is an open-source text-to-speech model based on an RVQ Transformer architecture with a Llama 3.2-style 8B backbone, designed fo… | 52 | 3224 | active |
| zai-org/GLM-4-Voice GLM-4-Voice is an end-to-end bilingual (Chinese/English) speech dialogue model from Zhipu AI, built on GLM-4-9B with a speech tokenizer and… | 22 | 3222 | active |
| BayLing-Models/BayLing-Speech LLaMA-Omni is an end-to-end speech interaction model built on Llama-3.1-8B-Instruct that generates simultaneous text and speech responses f… | 32 | 3146 | active |
| SuperCmdLabs/SuperCmd SuperCmd is an open-source macOS launcher combining Raycast-compatible extensions, hold-to-speak dictation, text-to-speech, AI chat and age… | 75 | 3124 | active |
| HeyWillow/willow Willow is an open source, self-hosted voice assistant platform for ESP32-S3-BOX hardware, designed as a privacy-focused alternative to Amaz… | 86 | 3097 | active |
| theajack/cnchar cnchar is a comprehensive TypeScript library for Chinese character processing, offering pinyin conversion, stroke counts, stroke order draw… | 66 | 3083 | active |
| elevenlabs/elevenlabs-python The official Python SDK for the ElevenLabs API, providing programmatic access to text-to-speech, speech-to-text, voice cloning, dubbing, mu… | 93 | 3078 | active |
| kyutai-labs/delayed-streams-modeling Kyutai's repository of Speech-To-Text and Text-To-Speech models built on the Delayed Streams Modeling framework, with implementations in Py… | 47 | 3017 | active |
| off-grid-ai/OGAM Off Grid AI (OGAM) is a cross-platform mobile and desktop application that runs AI entirely on-device: GGUF LLM chat with vision, Whisper s… | 78 | 3002 | active |
| KevinWang676/Bark-Voice-Cloning A one-click hub of Gradio Web UIs and Colab notebooks for open-source voice cloning, TTS, and voice conversion models including Bark, GPT-S… | 67 | 2947 | active |
| openai/openai-fm OpenAI.fm is an interactive web demo showcasing OpenAI's text-to-speech models, built with Next.js and the OpenAI Speech API. It lets users… | 51 | 2887 | active |
| Camb-ai/MARS5-TTS MARS5 is an open-source English text-to-speech model from CAMB.AI that uses a two-stage AR-NAR pipeline to generate expressive speech with … | 14 | 2817 | active |
| lnreader/lnreader LNReader is a free, open-source light novel and webnovel reader app for Android, inspired by Tachiyomi. It supports 200+ community-maintain… | 87 | 2793 | active |
| rhasspy/piper Piper is a fast, local neural text-to-speech system that runs offline on modest hardware, including Raspberry Pi devices. It offers many pr… | 10 | 11281 | maintenance |
| AutoArk/GPA GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an… | 54 | 2762 | active |
| hahahumble/speechgpt SpeechGPT is an open-source web application that lets users have voice conversations with ChatGPT using speech recognition and speech synth… | 62 | 2751 | active |
| quik-sms/quik QUIK is an open-source SMS messenger app for Android, a revived continuation of QKSMS. It replaces the stock messaging app with features li… | 87 | 2708 | active |
| FluidInference/FluidAudio A Swift SDK providing fully local, low-latency audio AI on Apple devices, including speech-to-text (Parakeet), speaker diarization (Pyannot… | 81 | 2699 | active |
| thewh1teagle/kokoro-onnx A Python library that runs the Kokoro text-to-speech model via ONNX Runtime, supporting CPU and GPU inference. It provides multi-language T… | 67 | 2680 | active |
| ZeframLou/call-me A minimal Claude Code plugin that places real phone calls to notify you when an AI coding agent finishes a task, gets stuck, or needs a dec… | 54 | 2638 | active |
| pndurette/gTTS gTTS is a Python library and CLI tool that interfaces with Google Translate's text-to-speech API to generate spoken MP3 audio from text. It… | 57 | 2627 | stable |
| nvaccess/nvda NVDA (NonVisual Desktop Access) is a free, open source screen reader for Microsoft Windows that reads on-screen text aloud via synthetic sp… | 90 | 2625 | stable |
| 6drf21e/ChatTTS_colab A one-click deployment wrapper around ChatTTS providing a Gradio web UI for text-to-speech, runnable in Google Colab or via an offline Wind… | 53 | 2592 | active |
| liou666/polyglot Polyglot is a cross-platform desktop (and web) application for practicing spoken language with AI conversation partners, built on ChatGPT f… | 39 | 2585 | active |
| ahmedeltaher/Android-MVVM-Architecture-Android-Voice-AI-SDK A reusable Android library (Kotlin, MVVM) that provides a full voice-driven AI conversation pipeline: microphone capture with VAD, speech-t… | 70 | 2578 | active |
| miantiao-me/hacker-podcast An AI-powered Chinese podcast application that automatically scrapes daily Hacker News top stories, generates Chinese summaries and scripts… | 62 | 2576 | active |
| zixiiu/Digital_Life_Server A Python server backend for a 'digital life' voice assistant that combines speech recognition, ChatGPT-based conversation, sentiment analys… | 30 | 2554 | active |
| AIGC-Audio/AudioGPT AudioGPT is a Python framework that wraps multiple audio foundation models (for speech, singing, sound, and talking-head tasks) behind a GP… | 30 | 10167 | maintenance |
| mozilla/TTS A deep learning library for advanced text-to-speech generation, built on PyTorch with models like Tacotron2, Glow-TTS, and various vocoders… | 23 | 10167 | maintenance |
| nateshmbhat/pyttsx3 pyttsx3 is an offline text-to-speech synthesis library for Python 3 that wraps system TTS engines like Sapi5, NSSpeechSynthesizer, and espe… | 67 | 2529 | active |
| janhq/ichigo Ichigo is a Python speech package for developers offering local realtime voice AI capabilities, including a compact 22M-parameter speech to… | 50 | 2492 | active |
| erew123/alltalk_tts AllTalk TTS is a text-to-speech application built on the Coqui TTS engine, usable standalone or as an extension for Text-generation-webui, … | 44 | 2429 | active |
| pnnbao97/VieNeu-TTS VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi… | 84 | 2427 | active |
| wan-h/awesome-digital-human-live2d An open-source digital human application that combines Live2D avatars with LLM-powered conversation, integrating ASR, LLM, TTS, and agent o… | 62 | 2413 | active |
| showlab/Paper2Video Paper2Video is a Python pipeline that automatically generates academic presentation videos from scientific papers, taking a paper PDF, a sp… | 48 | 2368 | active |
| jik876/hifi-gan The official PyTorch implementation of HiFi-GAN, a generative adversarial network that converts mel-spectrograms into high-fidelity 22.05 k… | 32 | 2367 | stable |
| Mentra-Community/MentraOS MentraOS is an open-source operating system and development platform for smart glasses, providing pairing, connection management, data stre… | 94 | 2326 | active |
| cosin2077/easyVoice EasyVoice is an open-source text-to-speech application that converts long texts and novels into high-quality audio with streaming playback … | 49 | 2287 | active |
| QwenAudio/qwen-audio-agent A realtime voice runtime and frontend for AI coding agents like Claude Code, Codex, and Qwen Code, letting agents talk, listen, and report … | 79 | 2264 | active |
| fishaudio/Bert-VITS2 Bert-VITS2 is a text-to-speech model implementation combining the VITS2 architecture with multilingual BERT embeddings, written in Python. … | 63 | 8796 | maintenance |
| rushindrasinha/youtube-shorts-pipeline Verticals v3 (repo youtube-shorts-pipeline) is a Python CLI that automates producing and publishing YouTube Shorts: it researches a topic, … | 61 | 2251 | active |
| lifeiteng/vall-e An unofficial PyTorch implementation of VALL-E, a zero-shot text-to-speech model that treats TTS as a conditional language modeling task ov… | 40 | 2215 | active |
| DigitalPhonetics/IMS-Toucan IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive… | 63 | 2207 | active |
| boson-ai/higgs-audio Higgs Audio is a text-audio foundation model project from Boson AI providing code and weights for conversational text-to-speech with zero-s… | 57 | 8329 | maintenance |
| MiniMax-AI/cli The official CLI for the MiniMax AI Platform, written in TypeScript, that generates text, images, video, speech, and music from the termina… | 81 | 2073 | active |
| kimjammer/Neuro A local recreation of the Neuro-Sama AI VTuber that runs open-source LLMs on consumer hardware, combining realtime speech-to-text, text-to-… | 28 | 2070 | active |
| travisvn/openai-edge-tts A self-hosted Python service that emulates the OpenAI text-to-speech API endpoint (/v1/audio/speech) using Microsoft Edge's free online TTS… | 26 | 2066 | active |
| Plachtaa/VALL-E-X An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint… | 10 | 7931 | maintenance |
| jaywalnut310/vits VITS is the official PyTorch implementation of an end-to-end text-to-speech model based on a conditional variational autoencoder with adver… | 32 | 7889 | maintenance |
| 0xShug0/audio.cpp audio.cpp is a pure C++ inference engine for audio models built on ggml, supporting TTS, STT, VAD, voice conversion, music generation, and … | 80 | 2022 | active |
| digitalsamba/claude-code-video-toolkit An AI-native video production toolkit designed for Claude Code, providing skills, commands, templates, and Python tools so an AI agent can … | 83 | 2008 | active |
| run-llama/notebookllama NotebookLlaMa is an open-source, Python-based alternative to Google's NotebookLM, backed by LlamaCloud for document ingestion and retrieval… | 52 | 1967 | active |
| praat/praat.github.io Praat is a desktop application for analyzing, synthesizing, and manipulating speech, widely used in phonetics research and teaching. It pro… | 99 | 1965 | stable |
| akdeb/ElatoAI ElatoAI is an open-source platform for running realtime voice AI conversations on ESP32/Arduino hardware, supporting 100+ STT, LLM, and TTS… | 64 | 1933 | active |
| AlexxIT/YandexStation A Home Assistant custom component (installed via HACS) for controlling Yandex smart speakers and other Alice smart home devices. It support… | 97 | 1917 | active |
| flybirdxx/ComfyUI-Qwen-TTS A ComfyUI custom node plugin that wraps Alibaba's Qwen3-TTS model for speech synthesis, zero-shot voice cloning, and natural-language voice… | 54 | 1876 | active |
| MixLabPro/comfyui-mixlab-nodes A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec… | 67 | 1863 | active |
| geometer/FBReaderJ FBReaderJ is the official repository behind FBReader, a long-running e-book reader application for Android written in Java. It supports for… | 32 | 1849 | active |
| RHVoice/RHVoice RHVoice is a free and open-source statistical parametric speech synthesizer (TTS) built on HTS technology, originally for Russian and now s… | 91 | 1831 | active |
| nazdridoy/kokoro-tts A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w… | 88 | 1816 | active |
| k2-fsa/sherpa-ncnn A C++ library for real-time offline speech recognition, text-to-speech, and voice activity detection built on the ncnn inference framework … | 58 | 1779 | active |
| wwbin2017/bailing Bailing is an open-source voice assistant application similar to GPT-4o, built with an ASR + VAD + LLM + TTS pipeline (FunASR, silero-vad, … | 54 | 1757 | active |
| High-Logic/Genie-TTS GENIE is a lightweight Python inference engine for the open-source GPT-SoVITS text-to-speech project, optimized for fast CPU-based speech s… | 61 | 1754 | active |
| ThioJoe/Auto-Synced-Translated-Dubs A Python CLI tool that automatically translates video subtitles into multiple languages and generates AI voice dubbed audio tracks synced t… | 59 | 1747 | active |
| ken107/read-aloud Read Aloud is a browser extension for Chrome, Firefox, and Edge that reads aloud webpage content using text-to-speech with one click. It wo… | 64 | 1731 | active |
| trymirai/uzu Uzu is a high-performance inference engine written in Rust for running AI models directly on-device, with Python, TypeScript, and Swift bin… | 85 | 1678 | active |
| LokerL/tts-vue TTS-Vue is a cross-platform desktop text-to-speech application built with Electron, Vue, ElementPlus, and Vite that uses Microsoft's Edge T… | 58 | 6101 | maintenance |
| mkiol/dsnote Speech Note is a Linux desktop and Sailfish OS application for note taking, reading, and translating text using offline Speech to Text, Tex… | 96 | 1607 | active |
| Norsico/Video-Materials-AutoGEN-Workstation A self-hosted short-video production workstation that combines AI script generation (Gemini), batch TTS voiceover, AI image asset synthesis… | 54 | 1599 | active |
| semperai/amica Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. … | 32 | 1593 | active |
| Alisa0808/vox-director An agent skill that turns a single topic into a finished Vox-style paper-collage explainer or ad video, automating script, keyframes, motio… | 56 | 1582 | active |
| Agents365-ai/video-podcast-maker An agent skill/plugin that turns a plain-language topic into a 4K narrated video podcast, combining research, script generation, multi-engi… | 79 | 1580 | active |
| MiniMax-AI/MiniMax-MCP The official MiniMax Model Context Protocol (MCP) server, written in Python, exposing MiniMax's text-to-speech, image generation, and video… | 64 | 1569 | active |
| hanmin0822/MisakaTranslator MisakaTranslator is a Windows desktop application that provides real-time machine translation for Galgames, text-based games, and manga. It… | 25 | 5746 | maintenance |
| Enemyx-net/VibeVoice-ComfyUI A ComfyUI custom node integration for Microsoft's VibeVoice text-to-speech model, providing single and multi-speaker voice synthesis with v… | 53 | 1549 | active |
| RunanywhereAI/RCLI RCLI is a single-binary CLI that runs open-source AI models locally on your machine, covering chat, vision, speech-to-text, text-to-speech,… | 82 | 1542 | active |
| vibevoice-community/VibeVoice VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker … | 60 | 1542 | active |
| dindin0497/SeeIt SeeIt is an inclusive Android app with two accessibility modes: one that converts spoken speech into text and plays corresponding ASL (Amer… | 41 | 1540 | active |
| AlekPet/ComfyUI_Custom_Nodes_AlekPet A collection of custom nodes for ComfyUI that extend its capabilities with painting, pose control, prompt translation, and speech recogniti… | 72 | 1524 | active |
| kyutai-labs/unmute Unmute is a system that lets any text LLM listen and speak by wrapping it with Kyutai's low-latency speech-to-text and text-to-speech model… | 60 | 1506 | active |
| met4citizen/TalkingHead A JavaScript library for real-time lip-synced talking avatars, rendering full-body 3D characters in the browser. It supports text-to-speech… | 67 | 1505 | active |
| stepfun-ai/Step-Audio2 Step-Audio 2 is an end-to-end multimodal large language model for industry-strength audio understanding and speech conversation, with open-… | 50 | 1503 | active |
| espressif/esp-sr ESP-SR is Espressif's speech recognition framework for ESP32-series chips, providing wake word detection (WakeNet), voice activity detectio… | 78 | 1492 | active |
| ekwek1/soprano Soprano is an ultra-lightweight 80M-parameter text-to-speech model and Python library for fast, expressive, high-fidelity speech synthesis … | 44 | 1486 | active |
| NVIDIA/tacotron2 NVIDIA's PyTorch implementation of the Tacotron 2 text-to-speech model, which synthesizes mel spectrograms from text for vocoder-based audi… | 32 | 5296 | maintenance |