domain: speech-processing
552 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Picovoice/porcupine Porcupine is a highly-accurate, lightweight on-device wake word detection engine powered by deep neural networks. It enables always-listeni… | 73 | 4922 | stable |
| buxuku/SmartSub SmartSub (妙幕) is a free, open-source cross-platform desktop application that provides an end-to-end subtitle and dubbing pipeline: speech-t… | 89 | 4778 | active |
| jianchang512/stt An offline local speech-to-text tool based on faster-whisper models that transcribes audio and video files into JSON, SRT subtitles, or pla… | 45 | 4766 | active |
| MoonshotAI/Kimi-Audio Kimi-Audio is an open-source audio foundation model (7B parameters) that unifies audio understanding, generation, and speech conversation i… | 31 | 4731 | active |
| WhisperSpeech/WhisperSpeech WhisperSpeech is an open-source text-to-speech system built by inverting OpenAI's Whisper model, aiming to be 'Stable Diffusion for speech'… | 56 | 4639 | active |
| Rikorose/DeepFilterNet DeepFilterNet is a low-complexity speech enhancement framework that performs real-time noise suppression on full-band 48kHz audio using dee… | 23 | 4632 | active |
| gradio-app/fastrtc FastRTC is a Python library that turns any Python function into a real-time audio and video stream over WebRTC or WebSockets. It includes b… | 58 | 4622 | active |
| fixie-ai/ultravox Ultravox is a fast multimodal LLM that understands human speech directly without a separate ASR stage, by projecting audio into the LLM's e… | 44 | 4555 | active |
| jing332/tts-server-android An Android system-wide text-to-speech (TTS) app that provides a TTS engine service with built-in Microsoft demo API support, custom HTTP TT… | 40 | 4483 | active |
| modelscope/ClearerVoice-Studio ClearerVoice-Studio is an open-source, AI-powered speech processing toolkit from ModelScope/Alibaba offering state-of-the-art pretrained mo… | 38 | 4445 | active |
| cmusphinx/pocketsphinx PocketSphinx is Carnegie Mellon's lightweight, speaker-independent continuous speech recognition engine, available as a C library, a comman… | 84 | 4335 | active |
| leetcode-mafia/cheetah Cheetah is a macOS app that provides real-time AI coaching for software engineering interviews. It transcribes conversation audio locally w… | 21 | 4257 | active |
| collabora/WhisperLive WhisperLive is a nearly-live speech-to-text application built on OpenAI's Whisper, serving real-time transcription over WebSocket and REST … | 86 | 4241 | active |
| huggingface/distil-whisper Distil-Whisper is a distilled version of OpenAI's Whisper model for English speech recognition, offering 6x faster inference, 49% fewer par… | 27 | 4112 | active |
| tmoroney/auto-subs AutoSubs is a local-first desktop application that generates AI subtitles on-device using Whisper, Moonshine, and Parakeet models, with spe… | 94 | 4091 | active |
| QwenLM/Qwen2.5-Omni Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre… | 31 | 4074 | active |
| MOSS-TTS MOSS-TTS-Nano is an open-source 0.1B-parameter multilingual speech generation (TTS) model from MOSI.AI and the OpenMOSS team, designed for … | 57 | 4031 | active |
| KoljaB/RealtimeTTS RealtimeTTS is a Python library that converts text, generators, and LLM token streams into speech audio with low latency. It supports multi… | 94 | 4016 | active |
| QwenLM/Qwen3-Omni Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an… | 52 | 3980 | active |
| digimata/quill Quill is an ultra-minimalist, fully local macOS meeting recorder and transcriber that runs as a single Swift menu-bar binary. It records mi… | 55 | 3874 | active |
| openai/openai-agents-js A lightweight, provider-agnostic TypeScript/JavaScript framework from OpenAI for building multi-agent LLM workflows, including text agents,… | 81 | 3714 | active |
| murtaza-nasir/speakr Speakr is a self-hosted web application that transcribes audio recordings and turns them into organized, searchable notes. It supports mult… | 84 | 3676 | active |
| kaldi-asr/kaldi Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea… | 52 | 15469 | maintenance |
| speaches-ai/speaches Speaches is an OpenAI API-compatible self-hosted server for speech-to-text (via faster-whisper), translation, and text-to-speech (via Kokor… | 78 | 3621 | active |
| facebookresearch/sam-audio SAM-Audio is Meta's foundation model for isolating any sound in audio using text, visual, or temporal prompts. This repository provides inf… | 55 | 3612 | active |
| SakiRinn/LiveCaptions-Translator A lightweight Windows application that combines the built-in Windows 11 LiveCaptions speech-to-text feature with translation APIs (includin… | 77 | 3585 | active |
| Soul-AILab/SoulX-Podcast SoulX-Podcast is the official inference codebase for a text-to-speech model that generates long-form, multi-turn, multi-speaker podcast-sty… | 42 | 3535 | active |
| neonbjb/tortoise-tts Tortoise TTS is a multi-voice text-to-speech library built on PyTorch that prioritizes highly realistic prosody and intonation. It combines… | 32 | 14870 | maintenance |
| common-voice/common-voice Mozilla Common Voice is a web platform for crowdsourcing voice donations to build public-domain speech datasets for training voice recognit… | 95 | 3484 | active |
| huangjunsen0406/py-xiaozhi py-xiaozhi is an open-source, cross-platform multimodal AI voice assistant client written in Python, compatible with the xiaozhi-esp32 ecos… | 85 | 3452 | active |
| Kedreamix/Linly-Talker Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, … | 48 | 3436 | active |
| Qwen3-ASR Qwen3-ASR is a family of open-source speech recognition models from Alibaba's Qwen team, supporting ASR and language identification across … | 55 | 3423 | active |
| WEIFENG2333/AsrTools AsrTools is a Python desktop application with a PyQt5-based GUI that converts audio and video files to text using online ASR engines, witho… | 48 | 3423 | active |
| libAudioFlux/audioFlux audioFlux is a C-based library with Python bindings for audio and music analysis and feature extraction. It supports dozens of time-frequen… | 53 | 3351 | active |
| xenova/whisper-web A browser-based speech recognition app that runs OpenAI's Whisper models entirely client-side using Transformers.js. It transcribes audio w… | 30 | 3338 | active |
| ahmetoner/whisper-asr-webservice A Dockerized REST webservice that wraps OpenAI Whisper (plus Faster Whisper and WhisperX engines) for automatic speech recognition. It expo… | 93 | 3326 | active |
| Open-Less/openless OpenLess is an open-source cross-platform voice input application that lets users hold a hotkey, speak, and have AI-polished text inserted … | 81 | 3321 | active |
| xiph/opus libopus is the reference implementation of the Opus audio codec, an IETF-standardized (RFC 6716) royalty-free codec for interactive speech … | 79 | 3293 | stable |
| rsxdalv/TTS-WebUI TTS WebUI is a free, open-source web interface combining Gradio and React that unifies 30+ AI models for text-to-speech, voice conversion, … | 89 | 3244 | active |
| echo-loop/Echo-Loop Echo Loop is an open-source Flutter-based English listening and speaking training app that guides learners through a structured listen-to-s… | 80 | 3242 | active |
| VOICEVOX VOICEVOX is a free, mid-quality text-to-speech (and singing synthesis) software whose editor is built with Electron, TypeScript, and Vue. I… | 93 | 3231 | active |
| MisoLabsAI/MisoTTS Miso TTS 8B is an open-source text-to-speech model based on an RVQ Transformer architecture with a Llama 3.2-style 8B backbone, designed fo… | 52 | 3224 | active |
| zai-org/GLM-4-Voice GLM-4-Voice is an end-to-end bilingual (Chinese/English) speech dialogue model from Zhipu AI, built on GLM-4-9B with a speech tokenizer and… | 22 | 3222 | active |
| wendy7756/AI-Video-Transcriber An open-source AI tool that transcribes, summarizes, and archives videos and podcasts from 30+ platforms (YouTube, TikTok, Bilibili, etc.) … | 62 | 3215 | active |
| Purfview/whisper-standalone-win Standalone Windows/Linux/macOS executables of OpenAI's Whisper and Faster-Whisper that transcribe audio and video to text and subtitles wit… | 50 | 3163 | active |
| BayLing-Models/BayLing-Speech LLaMA-Omni is an end-to-end speech interaction model built on Llama-3.1-8B-Instruct that generates simultaneous text and speech responses f… | 32 | 3146 | active |
| matthartman/ghost-pepper A free, open-source macOS menu bar app providing fully on-device speech-to-text dictation and meeting transcription using local Whisper, Pa… | 78 | 3140 | active |
| modelscope/3D-Speaker 3D-Speaker is an open-source Python toolkit for single- and multi-modal speaker verification, speaker recognition, and speaker diarization,… | 56 | 3121 | active |
| HeyWillow/willow Willow is an open source, self-hosted voice assistant platform for ESP32-S3-BOX hardware, designed as a privacy-focused alternative to Amaz… | 86 | 3097 | active |
| elevenlabs/elevenlabs-python The official Python SDK for the ElevenLabs API, providing programmatic access to text-to-speech, speech-to-text, voice cloning, dubbing, mu… | 93 | 3078 | active |
| kyutai-labs/delayed-streams-modeling Kyutai's repository of Speech-To-Text and Text-To-Speech models built on the Delayed Streams Modeling framework, with implementations in Py… | 47 | 3017 | active |
| CheshireCC/faster-whisper-GUI A desktop GUI application built with PySide6 for running faster-whisper and whisperX speech-to-text transcription. It lets users transcribe… | 71 | 2991 | active |
| KevinWang676/Bark-Voice-Cloning A one-click hub of Gradio Web UIs and Colab notebooks for open-source voice cloning, TTS, and voice conversion models including Bark, GPT-S… | 67 | 2947 | active |
| davabase/whisper_real_time A Python demo application that performs real-time speech-to-text transcription using OpenAI's Whisper model. It records audio continuously … | 39 | 2941 | active |
| facebookresearch/omnilingual-asr An open-source multilingual speech recognition library from Meta AI supporting over 1,600 languages, including hundreds never previously co… | 52 | 2898 | active |
| kitlangton/Hex Hex is a macOS application that converts your voice to text: press-and-hold a global hotkey to record, and it transcribes on-device and pas… | 85 | 2890 | active |
| openai/openai-fm OpenAI.fm is an interactive web demo showcasing OpenAI's text-to-speech models, built with Next.js and the OpenAI Speech API. It lets users… | 51 | 2887 | active |
| jhj0517/Whisper-WebUI A Gradio-based web interface for OpenAI's Whisper models that generates subtitles from files, YouTube videos, or microphone input. It suppo… | 63 | 2862 | active |
| linto-ai/whisper-timestamped A Python library extending OpenAI's Whisper models to produce accurate word-level timestamps and confidence scores during multilingual spee… | 79 | 2841 | active |
| Camb-ai/MARS5-TTS MARS5 is an open-source English text-to-speech model from CAMB.AI that uses a two-stage AR-NAR pipeline to generate expressive speech with … | 14 | 2817 | active |
| rhasspy/piper Piper is a fast, local neural text-to-speech system that runs offline on modest hardware, including Raspberry Pi devices. It offers many pr… | 10 | 11281 | maintenance |
| AutoArk/GPA GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an… | 54 | 2762 | active |
| hahahumble/speechgpt SpeechGPT is an open-source web application that lets users have voice conversations with ChatGPT using speech recognition and speech synth… | 62 | 2751 | active |
| rhasspy/rhasspy Rhasspy is a fully offline, privacy-focused set of voice assistant services supporting many human languages. It converts spoken voice comma… | 10 | 2750 | active |
| Starmel/OpenSuperWhisper OpenSuperWhisper is a macOS dictation application that records audio and transcribes it locally using Whisper or Parakeet models. It suppor… | 73 | 2721 | active |
| Vexa-ai/vexa Vexa is an open-source (Apache-2.0) meeting transcription API that dispatches bots to join Google Meet, Microsoft Teams, and Zoom calls and… | 87 | 2719 | active |
| dscripka/openWakeWord openWakeWord is an open-source Python library for detecting wake words (or phrases) in streaming audio, with pre-trained models for common … | 49 | 2702 | active |
| FluidInference/FluidAudio A Swift SDK providing fully local, low-latency audio AI on Apple devices, including speech-to-text (Parakeet), speaker diarization (Pyannot… | 81 | 2699 | active |
| thewh1teagle/kokoro-onnx A Python library that runs the Kokoro text-to-speech model via ONNX Runtime, supporting CPU and GPU inference. It provides multi-language T… | 67 | 2680 | active |
| Const-me/Whisper A Windows port of whisper.cpp that runs OpenAI's Whisper speech recognition model on the GPU via DirectCompute (Direct3D 11 compute shaders… | 60 | 10649 | maintenance |
| ZeframLou/call-me A minimal Claude Code plugin that places real phone calls to notify you when an AI coding agent finishes a task, gets stuck, or needs a dec… | 54 | 2638 | active |
| anliyuan/Ultralight-Digital-Human An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It… | 64 | 2627 | active |
| pndurette/gTTS gTTS is a Python library and CLI tool that interfaces with Google Translate's text-to-speech API to generate spoken MP3 audio from text. It… | 57 | 2627 | stable |
| 6drf21e/ChatTTS_colab A one-click deployment wrapper around ChatTTS providing a Gradio web UI for text-to-speech, runnable in Google Colab or via an offline Wind… | 53 | 2592 | active |
| asteroid-team/asteroid Asteroid is a PyTorch-based audio source separation toolkit for researchers, providing modular building blocks (filterbanks, encoders, mask… | 60 | 2584 | active |
| ahmedeltaher/Android-MVVM-Architecture-Android-Voice-AI-SDK A reusable Android library (Kotlin, MVVM) that provides a full voice-driven AI conversation pipeline: microphone capture with VAD, speech-t… | 70 | 2578 | active |
| zachlatta/freeflow FreeFlow is a free, open-source macOS dictation app that transcribes speech via Groq's fast API and pastes cleaned text into any focused te… | 77 | 2576 | active |
| zixiiu/Digital_Life_Server A Python server backend for a 'digital life' voice assistant that combines speech recognition, ChatGPT-based conversation, sentiment analys… | 30 | 2554 | active |
| AIGC-Audio/AudioGPT AudioGPT is a Python framework that wraps multiple audio foundation models (for speech, singing, sound, and talking-head tasks) behind a GP… | 30 | 10167 | maintenance |
| mozilla/TTS A deep learning library for advanced text-to-speech generation, built on PyTorch with models like Tacotron2, Glow-TTS, and various vocoders… | 23 | 10167 | maintenance |
| nateshmbhat/pyttsx3 pyttsx3 is an offline text-to-speech synthesis library for Python 3 that wraps system TTS engines like Sapi5, NSSpeechSynthesizer, and espe… | 67 | 2529 | active |
| microsoft/foundry-local Foundry Local is Microsoft's end-to-end local AI runtime and SDK suite (C#, JavaScript, Python, Rust) for running optimized models entirely… | 85 | 2525 | active |
| janhq/ichigo Ichigo is a Python speech package for developers offering local realtime voice AI capabilities, including a compact 22M-parameter speech to… | 50 | 2492 | active |
| Live-GalGame/LiveGalGame LiveGalGame is a playful app that overlays a visual-novel (galgame) interface onto real-life conversations, providing real-time speech-to-t… | 58 | 2490 | active |
| harry0703/AudioNotes AudioNotes is a locally-run web application that transcribes audio and video files into text and organizes them into structured Markdown no… | 67 | 2472 | active |
| erew123/alltalk_tts AllTalk TTS is a text-to-speech application built on the Coqui TTS engine, usable standalone or as an extension for Text-generation-webui, … | 44 | 2429 | active |
| pnnbao97/VieNeu-TTS VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi… | 84 | 2427 | active |
| yetone/voice-input-src A macOS menu-bar voice input app (Swift, macOS 14+) that lets users hold the Fn key to record and release to inject streaming-transcribed t… | 50 | 2409 | active |
| resemble-ai/resemble-enhance Resemble Enhance is an AI-powered Python tool that improves speech quality through denoising and enhancement, using a denoiser module and a… | 17 | 2397 | active |
| jik876/hifi-gan The official PyTorch implementation of HiFi-GAN, a generative adversarial network that converts mel-spectrograms into high-fidelity 22.05 k… | 32 | 2367 | stable |
| iver56/audiomentations Audiomentations is a Python library for audio data augmentation with an API inspired by albumentations. It provides fast CPU-based waveform… | 67 | 2314 | stable |
| cosin2077/easyVoice EasyVoice is an open-source text-to-speech application that converts long texts and novels into high-quality audio with streaming playback … | 49 | 2287 | active |
| yan5xu/ququ QuQu is an open-source, free desktop voice dictation app built with Electron, designed as a privacy-first Wispr Flow alternative optimized … | 31 | 2271 | active |
| QwenAudio/qwen-audio-agent A realtime voice runtime and frontend for AI coding agents like Claude Code, Codex, and Qwen Code, letting agents talk, listen, and report … | 79 | 2264 | active |
| fishaudio/Bert-VITS2 Bert-VITS2 is a text-to-speech model implementation combining the VITS2 architecture with multilingual BERT embeddings, written in Python. … | 63 | 8796 | maintenance |
| TEN-framework/ten-vad TEN VAD is a lightweight, low-latency voice activity detection library with an ONNX model for real-time speech detection. It supports C, Py… | 50 | 2248 | active |
| lifeiteng/vall-e An unofficial PyTorch implementation of VALL-E, a zero-shot text-to-speech model that treats TTS as a conditional language modeling task ov… | 40 | 2215 | active |
| DigitalPhonetics/IMS-Toucan IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive… | 63 | 2207 | active |
| TransWithAI/Faster-Whisper-TransWithAI-ChickenRice A high-performance audio/video transcription and translation application built on Faster Whisper, optimized for Japanese-to-Chinese transla… | 77 | 2196 | active |
| meizhong986/WhisperJAV WhisperJAV is a local, privacy-preserving subtitle generator for Japanese adult videos, combining Qwen3-ASR, Whisper, TEN-VAD/FireRedVAD vo… | 94 | 2170 | active |