domain: speech-processing
552 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| openai/whisper OpenAI's Whisper is a general-purpose speech recognition model and Python library built on a Transformer sequence-to-sequence architecture.… | 68 | 107981 | stable |
| RVC-Boss/GPT-SoVITS GPT-SoVITS is a Python-based few-shot voice cloning and text-to-speech system with an integrated WebUI. It supports zero-shot TTS from a 5-… | 68 | 61255 | active |
| microsoft/VibeVoice VibeVoice is Microsoft's open-source frontier voice AI framework combining a next-token diffusion text-to-speech model for expressive, long… | 60 | 53225 | active |
| ggml-org/whisper.cpp A high-performance C/C++ port of OpenAI's Whisper automatic speech recognition model with no external dependencies. It supports CPU, GPU (M… | 99 | 53203 | stable |
| jamiepine/voicebox Voicebox is a free, open-source, local-first AI voice studio desktop app that combines voice cloning, text-to-speech across 7 engines, and … | 75 | 51551 | active |
| 2noise/ChatTTS ChatTTS is a generative text-to-speech model optimized for conversational dialogue scenarios such as LLM assistants, supporting English and… | 69 | 39797 | active |
| myshell-ai/OpenVoice OpenVoice is a Python library and audio foundation model for instant voice cloning, requiring only a short reference audio clip to replicat… | 34 | 37312 | active |
| OpenBMB/VoxCPM VoxCPM is a tokenizer-free text-to-speech system built on a diffusion autoregressive architecture that generates continuous speech represen… | 79 | 36145 | active |
| fishaudio/fish-speech Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 l… | 74 | 32413 | active |
| cjpais/Handy Handy is a free, open-source, cross-platform desktop application for offline speech-to-text dictation. Press a configurable shortcut, speak… | 83 | 30402 | active |
| Zackriya-Solutions/meetily Meetily is a privacy-first, open-source AI meeting assistant that records, transcribes, and summarizes meetings entirely locally using Para… | 79 | 29940 | active |
| 78/xiaozhi-esp32 XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen… | 88 | 29186 | active |
| resemble-ai/chatterbox Chatterbox is a family of state-of-the-art open-source text-to-speech models by Resemble AI, including multilingual and low-latency Turbo v… | 52 | 26163 | active |
| mozilla-ai/llamafile llamafile is a Mozilla project that packages llama.cpp (and whisper.cpp for whisperfile) with Cosmopolitan Libc into single-file executable… | 94 | 25694 | active |
| SYSTRAN/faster-whisper A fast reimplementation of OpenAI's Whisper speech-to-text model built on the CTranslate2 inference engine, offering up to 4x speedup and l… | 57 | 25104 | active |
| m-bain/whisperX WhisperX is a Python library providing fast automatic speech recognition with word-level timestamps and speaker diarization. It combines Wh… | 91 | 23765 | active |
| index-tts/index-tts IndexTTS is an industrial-level zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest IndexTTS… | 78 | 23511 | active |
| QwenAudio/CosyVoice CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-… | 61 | 22925 | active |
| chidiwilliams/buzz Buzz is a cross-platform desktop application that transcribes and translates audio and video offline using OpenAI's Whisper models. It offe… | 95 | 21147 | active |
| DrewThomasson/ebook2audiobook A Python application that converts non-DRM e-books (e.g., EPUB, PDF) into audiobooks with chapters and metadata using TTS engines like XTTS… | 93 | 20053 | active |
| modelscope/FunASR FunASR is an industrial end-to-end speech recognition toolkit built on PyTorch, offering ASR, VAD, punctuation restoration, speaker diariza… | 99 | 20036 | active |
| nari-labs/dia Dia is a 1.6B-parameter open-weight text-to-speech model from Nari Labs that generates ultra-realistic multi-speaker dialogue in a single p… | 43 | 19378 | active |
| NVIDIA-NeMo/Speech NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, … | 98 | 18337 | active |
| KittenML/KittenTTS KittenTTS is an open-source, ultra-lightweight text-to-speech library built on ONNX, with models ranging from 15M to 80M parameters (25-80 … | 66 | 15403 | active |
| SWivid/F5-TTS F5-TTS is the official implementation of a fully non-autoregressive text-to-speech system based on flow matching with a Diffusion Transform… | 85 | 15167 | active |
| Vosk Vosk is an offline open source speech recognition toolkit supporting 20+ languages with small models and a streaming API. It provides bindi… | 66 | 15077 | stable |
| pipecat-ai/pipecat Pipecat is an open-source Python framework (BSD-2) for building real-time voice and multimodal conversational AI agents. It orchestrates sp… | 89 | 14769 | active |
| SesameAILabs/csm CSM (Conversational Speech Model) is Sesame's speech generation model that produces conversational audio from text and audio context, using… | 30 | 14720 | active |
| k2-fsa/sherpa-onnx sherpa-onnx is an offline speech processing toolkit built on next-gen Kaldi and onnxruntime, supporting speech-to-text, text-to-speech, spe… | 95 | 14411 | active |
| livekit/agents LiveKit Agents is an open-source framework for building realtime, multimodal voice AI agents that run as programmable participants in LiveK… | 90 | 13176 | active |
| QwenLM/Qwen3-TTS Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,… | 48 | 13113 | active |
| Vaibhavs10/insanely-fast-whisper A CLI tool for transcribing audio files on-device using OpenAI's Whisper models, powered by Hugging Face Transformers, Optimum, and Flash A… | 49 | 13049 | active |
| huggingface/speech-to-speech A low-latency, fully modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) built from open-source models, exposed via the OpenAI Realtime… | 85 | 12882 | active |
| PaddlePaddle/PaddleSpeech PaddleSpeech is an open-source speech and audio toolkit built on the PaddlePaddle deep learning platform, covering ASR with punctuation, st… | 66 | 12670 | active |
| abus-aikorea/voice-pro Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, … | 84 | 12647 | active |
| facebookresearch/seamless_communication A library of foundational multilingual multimodal AI models from Meta for speech and text translation, including SeamlessM4T, SeamlessExpre… | 71 | 11842 | active |
| rany2/edge-tts A Python module and CLI that lets you use Microsoft Edge's online neural text-to-speech service without needing Edge, Windows, or an API ke… | 79 | 11803 | active |
| speechbrain/speechbrain SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained… | 83 | 11785 | active |
| debpalash/VoiceStudio VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, vid… | 80 | 11737 | active |
| CorentinJ/Real-Time-Voice-Cloning A Python implementation of the SV2TTS (Transfer Learning from Speaker Verification to Multispeaker TTS) framework that clones a voice from … | 64 | 60110 | maintenance |
| TEN-framework/ten-framework TEN is an open-source framework for building real-time multimodal conversational AI agents, with a focus on low-latency voice assistants. I… | 85 | 11082 | active |
| SparkAudio/Spark-TTS Spark-TTS is an LLM-based text-to-speech system built on Qwen2.5 that generates speech directly from single-stream decoupled speech tokens.… | 27 | 11006 | active |
| QuentinFuxa/WhisperLiveKit WhisperLiveKit is a self-hosted, ultra-low-latency real-time speech-to-text pipeline built on state-of-the-art simultaneous speech research… | 87 | 10962 | active |
| kyutai-labs/moshi Moshi is a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai, built around the Mimi streaming neural audio… | 62 | 10949 | active |
| altic-dev/FluidVoice FluidVoice is an open-source (GPLv3) macOS dictation app that transcribes speech to text entirely on-device using models like Parakeet, Whi… | 84 | 10946 | active |
| moonshine-ai/moonshine Moonshine Voice is an open-source on-device AI toolkit providing very low latency speech-to-text, intent recognition, and text-to-speech fo… | 89 | 10944 | active |
| pyannote/pyannote-audio pyannote.audio is an open-source Python toolkit built on PyTorch for speaker diarization, providing neural building blocks like voice activ… | 95 | 10475 | active |
| xinnan-tech/xiaozhi-esp32-server A self-hosted backend server for the xiaozhi-esp32 open-source smart hardware project, implementing the Xiaozhi communication protocol over… | 81 | 10433 | active |
| NVIDIA/personaplex PersonaPlex is a real-time, full-duplex speech-to-speech conversational model from NVIDIA that supports persona control via text role promp… | 47 | 10393 | active |
| nexmoe/VidBee VidBee is a free, open-source Electron desktop app for downloading video and audio from 1000+ sites (via yt-dlp) and importing local media.… | 84 | 10386 | active |
| open-mmlab/Amphion Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s… | 51 | 10271 | active |
| KoljaB/RealtimeSTT RealtimeSTT is a Python library for low-latency, real-time speech-to-text with voice activity detection, wake word activation, and streamin… | 94 | 10080 | active |
| snakers4/silero-vad Silero VAD is a pre-trained, enterprise-grade Voice Activity Detector model available via PyPI, runnable with PyTorch or ONNX Runtime. It d… | 86 | 10061 | stable |
| espnet/espnet ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance… | 88 | 9941 | active |
| k2-fsa/OmniVoice OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style a… | 79 | 9455 | active |
| coqui-ai/TTS Coqui TTS is a deep learning toolkit for text-to-speech synthesis, providing pretrained models in over 1100 languages plus tools for traini… | 23 | 45953 | maintenance |
| QwenAudio/SenseVoice SenseVoice is an open-source speech foundation model (SenseVoiceSmall) providing multilingual ASR, spoken language identification, speech e… | 88 | 9151 | active |
| kyutai-labs/pocket-tts Pocket TTS is a lightweight 100M-parameter text-to-speech model and Python library from Kyutai that runs efficiently on CPUs without a GPU.… | 83 | 9135 | active |
| jianchang512/clone-voice A voice cloning tool with a web interface built on the coqui.ai xtts_v2 model, letting users synthesize speech in any voice from text or co… | 10 | 8989 | active |
| Uberi/speech_recognition A Python library for performing speech recognition with support for multiple engines and APIs, both online (Google, Azure, Wit.ai, OpenAI W… | 94 | 8987 | active |
| jasonppy/VoiceCraft VoiceCraft is a token infilling neural codec language model for zero-shot speech editing and text-to-speech on in-the-wild data like audiob… | 63 | 8574 | active |
| hexgrad/kokoro An inference library for Kokoro-82M, an open-weight text-to-speech model with 82 million parameters that delivers quality comparable to lar… | 36 | 8570 | active |
| netease-youdao/EmotiVoice EmotiVoice is an open-source text-to-speech engine supporting English and Chinese with over 2000 voices and prompt-controlled emotional syn… | 17 | 8523 | active |
| santinic/audiblez Audiblez is a Python CLI tool (with an optional GUI) that converts .epub e-books into .m4b audiobooks using the Kokoro-82M text-to-speech m… | 52 | 8460 | active |
| nl8590687/ASRT_SpeechRecognition ASRT is a deep-learning-based Chinese speech recognition (speech-to-text) system built with TensorFlow/Keras, using CNN, LSTM, attention me… | 57 | 8383 | active |
| GetStream/Vision-Agents An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O… | 84 | 8100 | active |
| suno-ai/bark Bark is Suno's open-source transformer-based text-to-audio model that generates highly realistic multilingual speech, music, background noi… | 30 | 39249 | maintenance |
| Blaizzy/mlx-audio MLX-Audio is a Python library built on Apple's MLX framework for fast text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS… | 88 | 7793 | active |
| jianchang512/ChatTTS-ui A local web interface for the ChatTTS text-to-speech model that synthesizes speech from mixed Chinese/English text, numbers, and symbols. I… | 70 | 7640 | active |
| babysor/MockingBird MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real… | 54 | 36909 | maintenance |
| myshell-ai/MeloTTS MeloTTS is a high-quality multi-lingual text-to-speech library supporting English (multiple accents), Spanish, French, Chinese, Japanese, a… | 16 | 7608 | active |
| Zyphra/Zonos Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I… | 24 | 7244 | active |
| thewh1teagle/vibe Vibe is a cross-platform desktop app for fully offline audio and video transcription using Whisper, Nemotron, and Parakeet models. It suppo… | 93 | 7206 | active |
| wzpan/wukong-robot wukong-robot is a modular Chinese-language voice assistant / smart speaker project in Python that combines offline wake-word detection, ASR… | 32 | 7126 | active |
| openai/openai-realtime-agents A Next.js TypeScript demo application showcasing advanced agentic patterns (Chat-Supervisor and Sequential Handoff) built on the OpenAI Rea… | 48 | 6966 | active |
| PaddlePaddle/models PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda… | 23 | 6932 | active |
| TalAter/annyang annyang is a tiny (2 KB), dependency-free JavaScript library that adds speech recognition and voice commands to websites using the Web Spee… | 78 | 6816 | stable |
| espeak-ng/espeak-ng eSpeak NG is a compact open-source text-to-speech synthesizer supporting over 100 languages and accents, using formant synthesis for small … | 67 | 6763 | active |
| HaujetZhao/CapsWriter-Offline CapsWriter-Offline is a fully offline voice input tool for Windows that transcribes speech to text when you hold CapsLock or mouse side but… | 92 | 6691 | active |
| argmaxinc/argmax-oss-swift A Swift SDK providing turn-key on-device speech AI frameworks for Apple Silicon, including WhisperKit (speech-to-text with Whisper), Speake… | 91 | 6338 | active |
| yl4579/StyleTTS2 StyleTTS 2 is a PyTorch text-to-speech model that uses style diffusion and adversarial training with large speech language models (e.g., Wa… | 29 | 6336 | active |
| canopyai/Orpheus-TTS Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a… | 45 | 6314 | active |
| neuphonic/neutts NeuTTS is a collection of open-source, on-device text-to-speech models built on small LLM backbones, with instant voice cloning from as lit… | 60 | 6256 | active |
| tyiannak/pyAudioAnalysis pyAudioAnalysis is a Python library for audio analysis covering feature extraction (MFCCs, spectrograms, chromagrams), supervised and unsup… | 48 | 6254 | active |
| modelscope/FunClip FunClip is an open-source, locally deployed video clipping tool that uses FunASR Paraformer models for speech recognition and subtitle gene… | 95 | 6190 | active |
| Beingpax/VoiceInk VoiceInk is a native macOS voice-to-text dictation app that transcribes speech to text almost instantly using local AI models (Parakeet, Wh… | 84 | 6115 | active |
| bytedance/MegaTTS3 MegaTTS 3 is ByteDance's open-source PyTorch text-to-speech model with a lightweight 0.45B-parameter Diffusion Transformer backbone. It pro… | 59 | 6091 | active |
| xiph/rnnoise RNNoise is a C library that uses a hybrid DSP/recurrent neural network approach for real-time full-band speech noise suppression. It also s… | 26 | 5801 | stable |
| OpenWhispr/openwhispr OpenWhispr is an open-source, privacy-first voice-to-text dictation desktop app for macOS, Windows, and Linux. It supports fully local tran… | 81 | 5766 | active |
| dnhkng/GLaDOS A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,… | 64 | 5689 | active |
| MahmoudAshraf97/whisper-diarization A pipeline that combines OpenAI Whisper transcription with speaker diarization using Voice Activity Detection (MarbleNet) and speaker embed… | 75 | 5630 | active |
| xiangyuecn/Recorder A JavaScript HTML5 audio recording library for browsers and hybrid apps, supporting mp3, wav, pcm, ogg, amr, webm, and g711 formats with re… | 84 | 5626 | active |
| huggingface/parler-tts Parler-TTS is a lightweight text-to-speech library from Hugging Face that generates high-quality, natural-sounding speech controllable via … | 25 | 5586 | active |
| dograh-hq/dograh Dograh is an open-source, self-hostable voice AI platform for building production voice agents, positioned as an alternative to Vapi and Re… | 84 | 5505 | active |
| modstart-lib/aigcpanel AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, … | 88 | 5482 | active |
| remsky/Kokoro-FastAPI A Dockerized FastAPI wrapper around the Kokoro-82M text-to-speech model exposing an OpenAI-compatible speech endpoint with CPU, NVIDIA, AMD… | 88 | 5373 | active |
| ysharma3501/LuxTTS LuxTTS is a lightweight zipvoice-based text-to-speech model for high-quality zero-shot voice cloning, generating clear 48kHz speech at up t… | 54 | 5306 | active |
| OHF-Voice/piper1-gpl Piper is a fast, fully local neural text-to-speech engine that embeds espeak-ng for phonemization and ships with a CLI, HTTP web server, Py… | 86 | 5287 | active |
| wenet-e2e/wenet WeNet is a production-first, end-to-end automatic speech recognition (ASR) toolkit built on PyTorch with transformer/conformer models. It p… | 62 | 5227 | active |
| Plachtaa/VITS-fast-fine-tuning A Python pipeline for fast fine-tuning of VITS text-to-speech models, enabling speaker adaptation in under an hour from short audio, long a… | 10 | 5012 | active |
page 1 / 6 next →