function: tts
498 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| harry0703/MoneyPrinterTurbo MoneyPrinterTurbo is an AI-powered application that generates high-definition short videos from a topic or keyword. It automatically writes… | 92 | 116915 | active |
| RVC-Boss/GPT-SoVITS GPT-SoVITS is a Python-based few-shot voice cloning and text-to-speech system with an integrated WebUI. It supports zero-shot TTS from a 5-… | 68 | 61255 | active |
| microsoft/VibeVoice VibeVoice is Microsoft's open-source frontier voice AI framework combining a next-token diffusion text-to-speech model for expressive, long… | 60 | 53225 | active |
| jamiepine/voicebox Voicebox is a free, open-source, local-first AI voice studio desktop app that combines voice cloning, text-to-speech across 7 engines, and … | 75 | 51551 | active |
| moeru-ai/airi Project AIRI is an open-source, self-hosted AI companion (waifu/virtual character) platform inspired by Neuro-sama, supporting Live2D and V… | 86 | 48451 | active |
| 2noise/ChatTTS ChatTTS is a generative text-to-speech model optimized for conversational dialogue scenarios such as LLM assistants, supporting English and… | 69 | 39797 | active |
| myshell-ai/OpenVoice OpenVoice is a Python library and audio foundation model for instant voice cloning, requiring only a short reference audio clip to replicat… | 34 | 37312 | active |
| OpenBMB/VoxCPM VoxCPM is a tokenizer-free text-to-speech system built on a diffusion autoregressive architecture that generates continuous speech represen… | 79 | 36145 | active |
| SillyTavern/SillyTavern SillyTavern is a locally installed, self-hosted web UI for interacting with text generation LLMs, image generation engines, and TTS voice m… | 87 | 32691 | active |
| fishaudio/fish-speech Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 l… | 74 | 32413 | active |
| koodo-reader/koodo-reader Koodo Reader is a cross-platform ebook manager and reader supporting EPUB, PDF, Kindle, comic archives, and many other formats. It offers c… | 94 | 27983 | active |
| ATH-MaaS/Pixelle-Video Pixelle-Video is an AI-powered fully automated short video engine that turns a single topic into a finished video. It automatically writes … | 67 | 27344 | active |
| resemble-ai/chatterbox Chatterbox is a family of state-of-the-art open-source text-to-speech models by Resemble AI, including multilingual and low-latency Turbo v… | 52 | 26163 | active |
| readest/readest Readest is an open-source, cross-platform ebook reader built with Next.js and Tauri v2, supporting EPUB, PDF, MOBI, AZW3, FB2, CBZ, and mor… | 84 | 23766 | active |
| index-tts/index-tts IndexTTS is an industrial-level zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest IndexTTS… | 78 | 23511 | active |
| QwenAudio/CosyVoice CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-… | 61 | 22925 | active |
| DrewThomasson/ebook2audiobook A Python application that converts non-DRM e-books (e.g., EPUB, PDF) into audiobooks with chapters and metadata using TTS engines like XTTS… | 93 | 20053 | active |
| nari-labs/dia Dia is a 1.6B-parameter open-weight text-to-speech model from Nari Labs that generates ultra-realistic multi-speaker dialogue in a single p… | 43 | 19378 | active |
| jianchang512/pyvideotrans pyVideoTrans is an open-source desktop application (with WebUI and CLI modes) that translates videos from one language to another. It provi… | 90 | 18812 | active |
| NVIDIA-NeMo/Speech NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, … | 98 | 18337 | active |
| Huanshere/VideoLingo VideoLingo is an all-in-one AI video translation, localization, and dubbing tool that produces Netflix-quality single-line subtitles with w… | 80 | 18266 | active |
| leon-ai/leon Leon is an open-source personal AI assistant that runs on your own server, combining tools, memory, and agentic execution with speech-to-te… | 86 | 17462 | active |
| KittenML/KittenTTS KittenTTS is an open-source, ultra-lightweight text-to-speech library built on ONNX, with models ranging from 15M to 80M parameters (25-80 … | 66 | 15403 | active |
| SWivid/F5-TTS F5-TTS is the official implementation of a fully non-autoregressive text-to-speech system based on flow matching with a Diffusion Transform… | 85 | 15167 | active |
| duixcom/Duix-Avatar Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea… | 57 | 14871 | active |
| SesameAILabs/csm CSM (Conversational Speech Model) is Sesame's speech generation model that produces conversational audio from text and audio context, using… | 30 | 14720 | active |
| k2-fsa/sherpa-onnx sherpa-onnx is an offline speech processing toolkit built on next-gen Kaldi and onnxruntime, supporting speech-to-text, text-to-speech, spe… | 95 | 14411 | active |
| tisfeng/Easydict Easydict is a concise and elegant macOS dictionary and translation app for looking up words and translating text, ready to use out of the b… | 95 | 14372 | stable |
| Open-LLM-VTuber/Open-LLM-VTuber Open-LLM-VTuber is a cross-platform AI companion application that enables real-time voice conversations with LLMs, featuring a Live2D avata… | 67 | 13473 | active |
| livekit/agents LiveKit Agents is an open-source framework for building realtime, multimodal voice AI agents that run as programmable participants in LiveK… | 90 | 13176 | active |
| QwenLM/Qwen3-TTS Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,… | 48 | 13113 | active |
| huggingface/speech-to-speech A low-latency, fully modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) built from open-source models, exposed via the OpenAI Realtime… | 85 | 12882 | active |
| PaddlePaddle/PaddleSpeech PaddleSpeech is an open-source speech and audio toolkit built on the PaddlePaddle deep learning platform, covering ASR with punctuation, st… | 66 | 12670 | active |
| abus-aikorea/voice-pro Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, … | 84 | 12647 | active |
| elebumm/RedditVideoMakerBot A Python application that automatically generates Reddit story videos for platforms like TikTok, YouTube, and Instagram with a single comma… | 72 | 12532 | active |
| YiiGuxing/TranslationPlugin A translation plugin for IntelliJ-based IDEs and Android Studio that supports multiple translation engines including Google, Microsoft, Dee… | 95 | 11831 | active |
| rany2/edge-tts A Python module and CLI that lets you use Microsoft Edge's online neural text-to-speech service without needing Edge, Windows, or an API ke… | 79 | 11803 | active |
| debpalash/VoiceStudio VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, vid… | 80 | 11737 | active |
| ankidroid/Anki-Android AnkiDroid is the Android client for the open-source Anki spaced repetition flashcard system, letting users study and sync decks with AnkiWe… | 93 | 11628 | active |
| CorentinJ/Real-Time-Voice-Cloning A Python implementation of the SV2TTS (Transfer Learning from Speaker Verification to Multispeaker TTS) framework that clones a voice from … | 64 | 60110 | maintenance |
| krillinai/KrillinAI KrillinAI is an AI-powered video translation and dubbing tool covering the full pipeline: video download, speech transcription, subtitle tr… | 84 | 11272 | active |
| SparkAudio/Spark-TTS Spark-TTS is an LLM-based text-to-speech system built on Qwen2.5 that generates speech directly from single-stream decoupled speech tokens.… | 27 | 11006 | active |
| kyutai-labs/moshi Moshi is a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai, built around the Mimi streaming neural audio… | 62 | 10949 | active |
| moonshine-ai/moonshine Moonshine Voice is an open-source on-device AI toolkit providing very low latency speech-to-text, intent recognition, and text-to-speech fo… | 89 | 10944 | active |
| linyqh/NarratoAI NarratoAI is an open-source, self-hosted AI-powered video commentary and automated editing tool. It uses large language models to generate … | 88 | 10879 | active |
| NVIDIA/personaplex PersonaPlex is a real-time, full-duplex speech-to-speech conversational model from NVIDIA that supports persona control via text role promp… | 47 | 10393 | active |
| niedev/RTranslator RTranslator is a free, open-source, offline real-time translation app for Android that runs speech recognition (Whisper) and translation (M… | 85 | 10355 | active |
| RunanywhereAI/runanywhere-sdks RunAnywhere is a set of cross-platform SDKs (Swift, Kotlin, React Native, Flutter, TypeScript, C++) over a shared C++ core for running AI m… | 85 | 10282 | active |
| open-mmlab/Amphion Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s… | 51 | 10271 | active |
| espnet/espnet ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance… | 88 | 9941 | active |
| ripperhe/Bob Bob is a macOS menu bar application for translation and OCR, supporting selection translation, screenshot translation, input translation, s… | 49 | 9735 | active |
| k2-fsa/OmniVoice OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style a… | 79 | 9455 | active |
| mengxi-ream/read-frog Read Frog is an open-source AI-powered browser extension for language learning and immersive translation. It translates web pages, selected… | 82 | 9346 | active |
| coqui-ai/TTS Coqui TTS is a deep learning toolkit for text-to-speech synthesis, providing pretrained models in over 1100 languages plus tools for traini… | 23 | 45953 | maintenance |
| kyutai-labs/pocket-tts Pocket TTS is a lightweight 100M-parameter text-to-speech model and Python library from Kyutai that runs efficiently on CPUs without a GPU.… | 83 | 9135 | active |
| jianchang512/clone-voice A voice cloning tool with a web interface built on the coqui.ai xtts_v2 model, letting users synthesize speech in any voice from text or co… | 10 | 8989 | active |
| jasonppy/VoiceCraft VoiceCraft is a token infilling neural codec language model for zero-shot speech editing and text-to-speech on in-the-wild data like audiob… | 63 | 8574 | active |
| hexgrad/kokoro An inference library for Kokoro-82M, an open-weight text-to-speech model with 82 million parameters that delivers quality comparable to lar… | 36 | 8570 | active |
| netease-youdao/EmotiVoice EmotiVoice is an open-source text-to-speech engine supporting English and Chinese with over 2000 voices and prompt-controlled emotional syn… | 17 | 8523 | active |
| santinic/audiblez Audiblez is a Python CLI tool (with an optional GUI) that converts .epub e-books into .m4b audiobooks using the Kokoro-82M text-to-speech m… | 52 | 8460 | active |
| duixcom/Duix-Mobile Duix Mobile is an open-source SDK for building real-time interactive AI avatars (digital humans) that run on-device on Android, iOS, tablet… | 69 | 8198 | active |
| GetStream/Vision-Agents An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O… | 84 | 8100 | active |
| suno-ai/bark Bark is Suno's open-source transformer-based text-to-audio model that generates highly realistic multilingual speech, music, background noi… | 30 | 39249 | maintenance |
| RayVentura/ShortGPT ShortGPT is a Python framework that automates AI-driven video content creation, combining LLM script generation, voiceover synthesis, asset… | 33 | 7884 | active |
| STranslate/STranslate STranslate is a ready-to-go Windows desktop translation and OCR tool built with WPF. It aggregates dozens of translation services (OpenAI, … | 95 | 7860 | active |
| Blaizzy/mlx-audio MLX-Audio is a Python library built on Apple's MLX framework for fast text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS… | 88 | 7793 | active |
| jianchang512/ChatTTS-ui A local web interface for the ChatTTS text-to-speech model that synthesizes speech from mixed Chinese/English text, numbers, and symbols. I… | 70 | 7640 | active |
| babysor/MockingBird MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real… | 54 | 36909 | maintenance |
| myshell-ai/MeloTTS MeloTTS is a high-quality multi-lingual text-to-speech library supporting English (multiple accents), Spanish, French, Chinese, Japanese, a… | 16 | 7608 | active |
| farzaa/clicky Clicky is an open-source macOS app that acts as an AI teacher living next to your cursor — it can see your screen, talk with you via voice,… | 50 | 7394 | active |
| Zyphra/Zonos Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I… | 24 | 7244 | active |
| wzpan/wukong-robot wukong-robot is a modular Chinese-language voice assistant / smart speaker project in Python that combines offline wake-word detection, ASR… | 32 | 7126 | active |
| ddean2009/MoneyPrinterPlus A Python desktop application that uses AI LLMs to batch-generate short videos with one click, including automatic video mashup/remixing and… | 30 | 7008 | active |
| yihong0618/xiaogpt A Python application that connects Xiaomi AI Speakers to ChatGPT and other LLMs, letting users ask questions starting with a trigger phrase… | 64 | 6901 | active |
| espeak-ng/espeak-ng eSpeak NG is a compact open-source text-to-speech synthesizer supporting over 100 languages and accents, using formant synthesis for small … | 67 | 6763 | active |
| microsoft/call-center-ai An AI-powered call center service that lets you initiate or receive phone calls handled by an LLM-driven agent via a simple API call. Built… | 64 | 6563 | active |
| souzatharsis/podcastfy Podcastfy is an open-source Python package and CLI that transforms multimodal content (websites, PDFs, images, YouTube videos, topics) into… | 60 | 6521 | active |
| argmaxinc/argmax-oss-swift A Swift SDK providing turn-key on-device speech AI frameworks for Apple Silicon, including WhisperKit (speech-to-text with Whisper), Speake… | 91 | 6338 | active |
| yl4579/StyleTTS2 StyleTTS 2 is a PyTorch text-to-speech model that uses style diffusion and adversarial training with large speech language models (e.g., Wa… | 29 | 6336 | active |
| canopyai/Orpheus-TTS Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a… | 45 | 6314 | active |
| neuphonic/neutts NeuTTS is a collection of open-source, on-device text-to-speech models built on small LLM backbones, with instant voice cloning from as lit… | 60 | 6256 | active |
| Shaunwei/RealChar RealChar is an open-source application for creating, customizing, and talking to AI characters/companions in realtime via voice or text. It… | 48 | 6213 | active |
| bytedance/MegaTTS3 MegaTTS 3 is ByteDance's open-source PyTorch text-to-speech model with a lightweight 0.45B-parameter Diffusion Transformer backbone. It pro… | 59 | 6091 | active |
| HapeLee/legado-with-MD3 Legado with MD3 is a free, open-source Android e-book and content aggregation reader, rebuilt from the Legado (阅读 3.0) project with a Mater… | 82 | 5860 | active |
| denizsafak/abogen Abogen is a desktop text-to-speech application that converts EPUB, PDF, text, markdown, and subtitle files into audiobooks with synchronize… | 75 | 5771 | active |
| dnhkng/GLaDOS A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,… | 64 | 5689 | active |
| huggingface/parler-tts Parler-TTS is a lightweight text-to-speech library from Hugging Face that generates high-quality, natural-sounding speech controllable via … | 25 | 5586 | active |
| dograh-hq/dograh Dograh is an open-source, self-hostable voice AI platform for building production voice agents, positioned as an alternative to Vapi and Re… | 84 | 5505 | active |
| remsky/Kokoro-FastAPI A Dockerized FastAPI wrapper around the Kokoro-82M text-to-speech model exposing an OpenAI-compatible speech endpoint with CPU, NVIDIA, AMD… | 88 | 5373 | active |
| liuzhao1225/YouDub-webui YouDub WebUI is an open-source AI video localization and dubbing tool that converts YouTube, Bilibili, or local videos into target-language… | 75 | 5358 | active |
| ysharma3501/LuxTTS LuxTTS is a lightweight zipvoice-based text-to-speech model for high-quality zero-shot voice cloning, generating clear 48kHz speech at up t… | 54 | 5306 | active |
| OHF-Voice/piper1-gpl Piper is a fast, fully local neural text-to-speech engine that embeds espeak-ng for phonemization and ships with a CLI, HTTP web server, Py… | 86 | 5287 | active |
| YILS-LIN/short-video-factory Short Video Factory is an open-source cross-platform desktop application that uses AI to generate and automatically edit short videos from … | 78 | 5176 | active |
| Plachtaa/VITS-fast-fine-tuning A Python pipeline for fast fine-tuning of VITS text-to-speech models, enabling speaker adaptation in under an hour from short audio, long a… | 10 | 5012 | active |
| buxuku/SmartSub SmartSub (妙幕) is a free, open-source cross-platform desktop application that provides an end-to-end subtitle and dubbing pipeline: speech-t… | 89 | 4778 | active |
| WhisperSpeech/WhisperSpeech WhisperSpeech is an open-source text-to-speech system built by inverting OpenAI's Whisper model, aiming to be 'Stable Diffusion for speech'… | 56 | 4639 | active |
| gradio-app/fastrtc FastRTC is a Python library that turns any Python function into a real-time audio and video stream over WebRTC or WebSockets. It includes b… | 58 | 4622 | active |
| jing332/tts-server-android An Android system-wide text-to-speech (TTS) app that provides a TTS engine service with built-in Microsoft demo API support, custom HTTP TT… | 40 | 4483 | active |
| pot-app/pot-desktop Pot is a cross-platform desktop application for hotkey-based text translation, screenshot OCR, and text-to-speech, built with Tauri. It sup… | 62 | 19343 | maintenance |
| dramaclaw/dramaclaw DramaClaw is a self-hosted, general-purpose AIGC video engine that turns scripts into finished films through a single pipeline covering sto… | 81 | 4355 | active |
page 1 / 5 next →