function: tts
498 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| LuckyHookin/edge-TTS-record A Windows desktop tool that records Microsoft Edge's online neural text-to-speech voices (e.g., Xiaoxiao, Yunyang) and saves the output as … | 23 | 1370 | maintenance |
| Jackywine/Bella Bella is a self-hosted Node.js web application that acts as a personalized AI digital companion with voice interaction. It combines Whisper… | 48 | 6379 | experimental |
| kripken/speak.js speak.js is a port of the eSpeak C++ speech synthesizer to JavaScript via Emscripten, enabling text-to-speech in the browser using only Jav… | 32 | 1335 | maintenance |
| kakaobrain/pororo PORORO is a Python library from Kakao Brain that provides unified access to dozens of pretrained neural models for natural language process… | 10 | 1305 | maintenance |
| sdkcarlos/artyom.js Artyom.js is a JavaScript library that wraps the Web Speech APIs (webkitSpeechRecognition and speechSynthesis) to add voice control, speech… | 23 | 1269 | maintenance |
| PlayVoice/vits_chinese A Chinese text-to-speech library combining VITS with BERT-based prosody embeddings and NaturalSpeech infer-loss features, supporting ONNX c… | 23 | 1227 | maintenance |
| bawangxx/XZVoice XZVoice is a free, open-source desktop text-to-speech application built with Electron, Vue, and ElementUI. It uses Alibaba Cloud's speech s… | 23 | 1168 | maintenance |
| spring-media/TransformerTTS A TensorFlow 2 implementation of a non-autoregressive Transformer-based neural network for text-to-speech synthesis, based on FastSpeech an… | 10 | 1161 | maintenance |
| Kyubyong/dc_tts A TensorFlow implementation of DC-TTS, a text-to-speech model based on deep convolutional networks with guided attention. It includes train… | 32 | 1156 | maintenance |
| openinterpreter/01 An open-source voice interface platform that lets users control computers conversationally, powered by Open Interpreter. It pairs a Python … | 16 | 5156 | experimental |
| RafalWilinski/telegram-chatgpt-concierge-bot A self-hosted Telegram bot that lets you chat with OpenAI's ChatGPT via text and voice messages. It uses LangChainJS for conversation histo… | 30 | 1130 | maintenance |
| synesthesiam/opentts OpenTTS is a text-to-speech server that unifies access to multiple open source TTS systems (Larynx, Glow-Speak, Coqui-TTS, MaryTTS, flite, … | 10 | 1118 | maintenance |
| Syan-Lin/CyberWaifu CyberWaifu is a Python chatbot application that combines LLMs (ChatGPT, Claude) with TTS (edge-tts, Azure) to create realistic conversation… | 30 | 1099 | maintenance |
| JiehangXie/PaddleBoBo PaddleBoBo is a Python project built on PaddlePaddle (with PaddleSpeech and PaddleGAN) that quickly generates a virtual streamer (VTuber) f… | 32 | 1062 | maintenance |
| Yink/Amadeus An Android app that replicates the Amadeus AI assistant app from Steins;Gate 0, built primarily for cosplay purposes. It features speech re… | 23 | 1061 | maintenance |
| Edresson/YourTTS YourTTS is a zero-shot multi-speaker text-to-speech and voice conversion model built on VITS, implemented in the Coqui TTS framework. It su… | 23 | 1053 | maintenance |
| bravekingzhang/text2video A Python web application that converts text (e.g., novel passages) into narrated videos. It splits text into sentences, generates images vi… | 29 | 1049 | maintenance |
| cbh123/narrator A Python app that watches your webcam and generates David Attenborough-style narration of what it sees, using GPT vision models and ElevenL… | 69 | 4426 | experimental |
| susiai/susi_device SUSI Device provides sources to install the SUSI AI assistant stack on a Raspberry Pi, combining microphone/speaker, a small display, a loc… | 32 | 1004 | maintenance |
| NATSpeech/NATSpeech A PyTorch framework for non-autoregressive text-to-speech (NAR-TTS), containing official implementations of PortaSpeech (NeurIPS 2021) and … | 23 | 1004 | maintenance |
| enhuiz/vall-e An unofficial PyTorch implementation of the VALL-E text-to-speech audio language model, built on the EnCodec tokenizer. It provides trainin… | 31 | 2976 | experimental |
| Priler/jarvis JARVIS is an offline, privacy-respecting voice assistant built in Rust with Tauri, using neural networks for speech-to-text, text-to-speech… | 53 | 2909 | experimental |
| yakami129/VirtualWife VirtualWife is a self-hosted virtual digital human (AI companion) application that combines VRM 3D character models with LLM-powered conver… | 20 | 2889 | experimental |
| fikrikarim/parlor Parlor is a fully on-device, real-time multimodal voice assistant similar to GPT-Live, combining speech recognition, a Gemma vision-languag… | 78 | 2039 | experimental |
| Mini-Omni Mini-Omni is an open-source multimodal large language model that performs real-time end-to-end speech-to-speech conversation with streaming… | 22 | 1922 | experimental |
| collabora/WhisperFusion WhisperFusion is a real-time voice chat application that combines WhisperLive speech-to-text, a Mistral/Phi LLM, and WhisperSpeech text-to-… | 26 | 1647 | experimental |
| lucidrains/naturalspeech2-pytorch A PyTorch implementation of NaturalSpeech 2, a zero-shot text-to-speech and singing synthesizer that combines a neural audio codec with a l… | 20 | 1333 | experimental |
| elfvingralf/macOSpilot-ai-assistant macOSpilot is a macOS desktop AI assistant built with Electron that answers spoken or typed questions about whatever application is current… | 26 | 1157 | experimental |
| supertone-inc/supertonic Supertonic is a lightning-fast, on-device multilingual text-to-speech system powered by ONNX Runtime, with a compact 99M-parameter open-wei… | 58 | 13734 | abandoned |
| idootop/mi-gpt MiGPT is a Node.js/Docker application that connects Xiaomi's XiaoAI smart speakers to ChatGPT and Doubao, turning them into custom AI voice… | 10 | 12508 | abandoned |
| idootop/open-xiaoai Open-XiaoAI is a Rust-based client/server project that takes over the audio input and output of Xiaomi XiaoAI smart speakers (LX06 and OH2P… | 10 | 2595 | abandoned |
| C-Nedelcu/talk-to-chatgpt A Chrome and Edge browser extension that lets users talk to ChatGPT using speech recognition and hear responses via text-to-speech, with op… | 32 | 1929 | abandoned |
| wzpan/dingdang-robot Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi. It uses pluggable STT/TTS en… | 10 | 1873 | abandoned |
| justLV/onju-voice A hackable AI home assistant platform that replaces the internals of a Google Nest Mini with a custom ESP32-S3 PCB, paired with a server th… | 63 | 1711 | abandoned |
| elevenlabs/elevenlabs-mcp The official ElevenLabs Model Context Protocol (MCP) server, exposing ElevenLabs text-to-speech, speech-to-text, voice design, and conversa… | 10 | 1532 | abandoned |
| fossasia/susi_smart_box SUSI.AI Smart Box is an open-source smart speaker / voice assistant hardware project built around the SUSI.AI assistant. It provides the so… | 10 | 1529 | abandoned |
| idootop/migpt-next MiGPT-Next is a TypeScript library and Docker-runnable service that connects Xiaomi's XiaoAI smart speakers to OpenAI-compatible large lang… | 10 | 1442 | abandoned |
| zenorocha/voice-elements A pair of Polymer-based Web Components (<voice-player> and <voice-recognition>) that wrap the Web Speech API for speech synthesis (text to … | 23 | 1348 | abandoned |
| dingdang-robot/dingdang-robot Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi and other Linux hosts. It use… | 10 | 1330 | abandoned |
| MycroftAI/mimic3 Mimic 3 is a fast, local neural text-to-speech engine developed by Mycroft for the Mark II voice assistant, usable as a Python library, CLI… | 29 | 1264 | abandoned |
| BrasD99/HeyGenClone An open-source Python application that clones the HeyGen video translation system, translating videos into multiple languages with voice ov… | 10 | 1027 | abandoned |
| Unsloth Unsloth is a desktop application for running and fine-tuning LLMs, diffusion, embedding, and audio models locally, with support for NVIDIA,… | 94 | 74883 | active |
| calesthio/OpenMontage OpenMontage is an open-source agentic video production system that turns AI coding assistants (Claude Code, Codex, Cursor, Copilot) into fu… | 59 | 51272 | active |
| mudler/LocalAI LocalAI is an open-source, self-hosted AI inference runtime that runs LLMs, vision, speech, image, and video models behind OpenAI/Anthropic… | 93 | 48696 | active |
| lfnovo/open-notebook Open Notebook is an open-source, privacy-focused alternative to Google's Notebook LM that lets users organize multi-modal sources (PDFs, vi… | 86 | 37711 | active |
| 78/xiaozhi-esp32 XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen… | 88 | 29186 | active |
| openai/openai-agents-python OpenAI Agents SDK is a lightweight Python framework for building multi-agent LLM workflows with agents, handoffs, tools, guardrails, sessio… | 82 | 28982 | active |
| facebookresearch/audiocraft A PyTorch library from Meta for audio processing and generation with deep learning, featuring the EnCodec neural audio codec and generative… | 61 | 23586 | active |
| livekit/livekit LiveKit is an open-source, scalable WebRTC SFU media server written in Go that provides realtime video, audio, and data transport for appli… | 99 | 20530 | stable |
| huggingface/transformers.js Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) … | 91 | 16270 | active |
| alibaba/MNN MNN is a lightweight, high-performance deep learning inference engine developed by Alibaba, supporting LLMs, vision, audio, and multimodal … | 93 | 15973 | active |
| pipecat-ai/pipecat Pipecat is an open-source Python framework (BSD-2) for building real-time voice and multimodal conversational AI agents. It orchestrates sp… | 89 | 14769 | active |
| SubtitleEdit/subtitleedit Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac… | 94 | 13971 | active |
| FujiwaraChoki/MoneyPrinter MoneyPrinter is a Python application that automates the creation of YouTube Shorts from a user-provided video topic, using local Ollama mod… | 59 | 13891 | active |
| xszyou/Fay Fay is an open-source Python digital human framework that connects 2.5D/3D/mobile/web digital humans or OpenAI-compatible LLMs to business … | 75 | 13455 | active |
| ace-step/ACE-Step-1.5 ACE-Step 1.5 is an open-source music generation foundation model combining a language model planner with a Diffusion Transformer to create … | 79 | 12421 | active |
| h2oai/h2ogpt h2oGPT is an Apache-2.0 open-source application for chatting with local, private LLMs and querying/summarizing your own documents (PDFs, Wo… | 10 | 11969 | active |
| speechbrain/speechbrain SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained… | 83 | 11785 | active |
| Baiyuetribe/paper2gui Paper2GUI (now branded 小白兔AI / Xiaobaitu AI) is a desktop AI toolbox application that packages 50+ AI research models into install-free GUI… | 23 | 10709 | active |
| VoltAgent/voltagent VoltAgent is an open-source TypeScript framework for building production-ready AI agents with memory, tools, RAG, guardrails, MCP, voice, a… | 80 | 10425 | active |
| lipku/LiveTalking LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit… | 89 | 9238 | active |
| fonoster/fonoster Fonoster is an open-source programmable telecommunications stack (a Twilio alternative) for building voice and messaging applications, with… | 99 | 8080 | active |
| a-ghorbani/pocketpal-ai PocketPal AI is an open-source mobile app that runs GGUF language models entirely on-device for private, offline chat. It supports model do… | 85 | 8067 | active |
| enricoros/big-AGI Big-AGI is an open-source AI workspace web application that lets users chat with many state-of-the-art LLM providers (OpenAI, Anthropic, Ge… | 93 | 7102 | active |
| PaddlePaddle/models PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda… | 23 | 6932 | active |
| gaozhangmin/boxplayer BoxPlayer is a cross-platform desktop application that unifies multiple cloud drives (Aliyun Drive, Baidu, 115, Quark, OneDrive, etc.), loc… | 97 | 6878 | active |
| Dooy/chatgpt-web-midjourney-proxy A unified web/desktop UI for ChatGPT plus AI image, music, and video generation services like Midjourney, Suno, Luma, Runway, and Flux. It … | 76 | 6785 | active |
| multimodal-art-projection/YuE YuE is a family of open-source foundation models based on the LLaMA2 architecture that generate full songs (up to five minutes) with vocals… | 32 | 6403 | active |
| vllm-project/vllm-omni vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer… | 83 | 6369 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| modstart-lib/aigcpanel AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, … | 88 | 5482 | active |
| lemonade-sdk/lemonade Lemonade is a local AI server that runs optimized LLMs (plus image, speech, and embedding models) on your own GPU and NPU, exposing OpenAI-… | 82 | 5472 | active |
| signalwire/freeswitch FreeSWITCH is an open-source software-defined telecom stack that turns commodity servers into a full telephony platform for voice, video, a… | 94 | 5118 | stable |
| MoonshotAI/Kimi-Audio Kimi-Audio is an open-source audio foundation model (7B parameters) that unifies audio understanding, generation, and speech conversation i… | 31 | 4731 | active |
| Osmantic/ODS ODS (Osmantic Deployment System) is a self-hosted AI server installer and runtime that turns a PC, Mac, or Linux machine into a private AI … | 80 | 4716 | active |
| ArcReel/ArcReel ArcReel is an open-source, self-hosted AI video production workspace that turns novels, scripts, or product material into characters, scene… | 82 | 4211 | active |
| QwenLM/Qwen2.5-Omni Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre… | 31 | 4074 | active |
| OpenMinis/OpenMinis OpenMinis is a free, open-source mobile AI agent app for iOS and Android that connects leading LLM providers (Claude, GPT, Gemini, etc.) vi… | 69 | 3985 | active |
| claraverse-space/ClaraVerse ClaraVerse is a self-hosted, privacy-focused AI workspace that combines chat, multi-agent teams, a visual workflow builder, RAG document pr… | 80 | 3894 | active |
| zly2006/zhihu-plus-plus Zhihu++ is an open-source third-party Android client for the Chinese Q&A platform Zhihu, built in Kotlin, that removes ads, promotional pos… | 85 | 3845 | active |
| tbxark/ChatGPT-Telegram-Workers A Telegram ChatGPT bot deployable on Cloudflare Workers, Vercel, or Docker with a single-file, dependency-free setup. It supports multiple … | 66 | 3811 | active |
| Chevey339/kelivo Kelivo is an open-source, cross-platform LLM chat client built with Flutter that runs on Android, iOS, HarmonyOS, Windows, macOS, and Linux… | 84 | 3766 | active |
| OvidijusParsiunas/deep-chat Deep Chat is a fully customizable, framework-agnostic AI chatbot web component that can be added to any website with one line of code. It s… | 91 | 3704 | active |
| sukeesh/Jarvis Jarvis is a command-line personal assistant for Linux, macOS, and Windows written in Python. It offers 15+ task categories including weathe… | 48 | 3645 | active |
| sligter/LandPPT LandPPT is an AI-powered presentation generation platform that turns a topic or uploaded documents (PDF, Word, Markdown, Excel, PPT) into p… | 81 | 3575 | active |
| n3d1117/chatgpt-telegram-bot A self-hosted Telegram bot written in Python that integrates with OpenAI's official ChatGPT, DALL·E, and Whisper APIs to answer questions, … | 33 | 3460 | active |
| Grt1228/chatgpt-java An unofficial Java SDK for the OpenAI API covering all official endpoints including chat completions (GPT-3.5/GPT-4), DALL-E image generati… | 21 | 3423 | active |
| chenyme/Chenyme-AAVT Chenyme-AAVT is a fully automated audio/video translation application that uses Whisper (faster-whisper) for speech recognition, large lang… | 10 | 3128 | active |
| Snailclimb/interview-guide An open-source AI-powered interview platform built with Spring Boot 4.1, Java 25, Spring AI 2.0, React, PostgreSQL/pgvector, and Redis. It … | 59 | 3106 | active |
| TanStack/ai TanStack AI is a type-safe, provider-agnostic TypeScript SDK for building AI applications with streaming chat, tool calling, agents, struct… | 80 | 3028 | active |
| SharpAI/DeepCamera DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r… | 86 | 3019 | active |
| MacPaw/OpenAI A community-maintained Swift package that wraps the OpenAI public API, supporting chat completions, responses, function calling, MCP tools,… | 94 | 2935 | active |
| OpenMind/OM1 OM1 is a modular AI runtime and hardware abstraction layer for building multimodal AI agents that run on physical robots and in simulators.… | 82 | 2897 | active |
| datascale-ai/opentalking OpenTalking is an open-source Python framework for building real-time AI digital-human (talking avatar) conversation products. It orchestra… | 65 | 2897 | active |
| luoluoluo22/jianying-editor-skill A Python-based AI agent skill that automates video editing in JianYing (the Chinese desktop version of CapCut) by directly generating and i… | 55 | 2878 | active |
| cdhigh/KindleEar KindleEar is a self-hostable Python web application that aggregates RSS/ATOM/JSON feeds and web content (including Calibre recipes) into ep… | 73 | 2865 | active |
| sahibzada-allahyar/YC-Killer A collection of open-source, enterprise-grade AI agents intended as free alternatives to commercial Y Combinator startups. It includes a de… | 64 | 2792 | active |
| ClassIsland/ClassIsland ClassIsland is a cross-platform desktop application that displays class schedules and related information on classroom multimedia screens, … | 92 | 2709 | active |
| Vali-98/ChatterUI ChatterUI is a native mobile frontend for running large language models on-device via llama.cpp or connecting to commercial and open-source… | 83 | 2689 | active |
| Project-N-E-K-O/N.E.K.O Project N.E.K.O. is an open-source AI companion application — a proactive catgirl-style AI that lives on your desktop, initiates interactio… | 85 | 2681 | active |