domain: speech-processing
552 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Nutlope/notesGPT NotesGPT is an open-source AI-powered voice note-taking web app that records voice notes, transcribes them with Whisper, and generates summ… | 68 | 2153 | active |
| boson-ai/higgs-audio Higgs Audio is a text-audio foundation model project from Boson AI providing code and weights for conversational text-to-speech with zero-s… | 57 | 8329 | maintenance |
| kaixxx/noScribe noScribe is a free, open-source desktop application that transcribes audio locally using OpenAI's Whisper (via faster-whisper) and pyannote… | 78 | 2132 | active |
| kleinlee/DH_live DH_live (mini) is an open-source 2D talking-head digital human toolkit that generates real-time lip-synced avatar video from a single refer… | 67 | 2131 | active |
| QwenLM/Qwen2-Audio Official repository for Qwen2-Audio, a 7B-parameter large audio-language model from Alibaba Cloud that accepts audio inputs and responds to… | 31 | 2099 | active |
| schibsted/WAAS Whisper as a Service (WAAS) is a self-hosted GUI and API wrapper around OpenAI's Whisper speech-to-text model, with asynchronous job queuin… | 73 | 2075 | active |
| kimjammer/Neuro A local recreation of the Neuro-Sama AI VTuber that runs open-source LLMs on consumer hardware, combining realtime speech-to-text, text-to-… | 28 | 2070 | active |
| travisvn/openai-edge-tts A self-hosted Python service that emulates the OpenAI text-to-speech API endpoint (/v1/audio/speech) using Microsoft Edge's free online TTS… | 26 | 2066 | active |
| Plachtaa/VALL-E-X An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint… | 10 | 7931 | maintenance |
| jaywalnut310/vits VITS is the official PyTorch implementation of an end-to-end text-to-speech model based on a conditional variational autoencoder with adver… | 32 | 7889 | maintenance |
| ricky0123/vad A JavaScript/TypeScript library that runs Silero VAD via ONNX Runtime Web to detect voice activity directly in the browser. It provides a s… | 69 | 2043 | active |
| mli/autocut AutoCut is a Python CLI tool that automatically transcribes video audio into subtitles using Whisper, then cuts video segments based on whi… | 32 | 7790 | maintenance |
| 0xShug0/audio.cpp audio.cpp is a pure C++ inference engine for audio models built on ggml, supporting TTS, STT, VAD, voice conversion, music generation, and … | 80 | 2022 | active |
| juanmc2005/diart Diart is a Python framework for building AI-powered real-time audio applications, best known for state-of-the-art streaming speaker diariza… | 62 | 2022 | active |
| FireRedTeam/FireRedASR FireRedASR is a family of open-source industrial-grade automatic speech recognition models supporting Mandarin, Chinese dialects, and Engli… | 52 | 1971 | active |
| praat/praat.github.io Praat is a desktop application for analyzing, synthesizing, and manipulating speech, widely used in phonetics research and teaching. It pro… | 99 | 1965 | stable |
| kadirnar/whisper-plus A Python library wrapping OpenAI Whisper-family models (including distil-whisper and MLX variants) for fast speech-to-text transcription wi… | 60 | 1956 | active |
| akdeb/ElatoAI ElatoAI is an open-source platform for running realtime voice AI conversations on ESP32/Arduino hardware, supporting 100+ STT, LLM, and TTS… | 64 | 1933 | active |
| LCAV/pyroomacoustics Pyroomacoustics is a Python package for audio signal processing in indoor scenarios, combining a fast C++ room acoustics simulator (image s… | 90 | 1928 | stable |
| bugbakery/audapolis Audapolis is a free, open-source desktop editor for spoken-word audio and video that transcribes speech automatically and lets you edit med… | 69 | 1890 | active |
| flybirdxx/ComfyUI-Qwen-TTS A ComfyUI custom node plugin that wraps Alibaba's Qwen3-TTS model for speech synthesis, zero-shot voice cloning, and natural-language voice… | 54 | 1876 | active |
| timsainb/noisereduce A Python library for noise reduction in time-domain signals using spectral gating, supporting stationary and non-stationary noise reduction… | 39 | 1874 | active |
| lanbinleo/bili2text bili2text is a Python command-line tool that converts Bilibili videos into text transcripts from a link or BV number, handling download, au… | 56 | 1873 | active |
| MontrealCorpusTools/Montreal-Forced-Aligner Montreal Forced Aligner is a command line utility for time-aligning orthographic transcriptions and pronunciation dictionary entries to aud… | 98 | 1872 | active |
| dindin0497/HearIt HearIt is an Android accessibility app that captures microphone audio and transcribes it to text in real time, displaying large on-screen c… | 62 | 1842 | active |
| handy-computer/transcribe.cpp A C/C++ speech-to-text inference library built on the ggml runtime that runs 16+ ASR model families (Whisper, Parakeet, Canary, Moonshine, … | 81 | 1836 | active |
| RHVoice/RHVoice RHVoice is a free and open-source statistical parametric speech synthesizer (TTS) built on HTS technology, originally for Russian and now s… | 91 | 1831 | active |
| nazdridoy/kokoro-tts A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w… | 88 | 1816 | active |
| FL33TW00D/whisper-turbo Whisper Turbo is a fast, cross-platform, GPU-accelerated implementation of OpenAI's Whisper speech recognition model, built on the Ratchet … | 19 | 1794 | active |
| abb128/LiveCaptions LiveCaptions is a Linux desktop application that displays real-time captions for desktop or microphone audio using a local speech recogniti… | 26 | 1783 | active |
| k2-fsa/sherpa-ncnn A C++ library for real-time offline speech recognition, text-to-speech, and voice activity detection built on the ncnn inference framework … | 58 | 1779 | active |
| wwbin2017/bailing Bailing is an open-source voice assistant application similar to GPT-4o, built with an ASR + VAD + LLM + TTS pipeline (FunASR, silero-vad, … | 54 | 1757 | active |
| High-Logic/Genie-TTS GENIE is a lightweight Python inference engine for the open-source GPT-SoVITS text-to-speech project, optimized for fast CPU-based speech s… | 61 | 1754 | active |
| x007xyz/flycut-caption FlyCut Caption is an AI-powered video subtitle editing tool built as a React component and web/desktop app, offering speech recognition wit… | 78 | 1753 | active |
| TypeWhisper/typewhisper-mac TypeWhisper is a macOS application for local, on-device speech-to-text dictation built on WhisperKit, with optional cloud transcription via… | 78 | 1735 | active |
| AutoArk/EVA-OS EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining … | 68 | 1712 | active |
| OpenMOSS/MOSS-Transcribe-Diarize MOSS-Transcribe-Diarize 0.9B is an open-source end-to-end audio understanding model that jointly performs multi-speaker speech transcriptio… | 58 | 1695 | active |
| absadiki/subsai Subs AI is a subtitles generation tool available as a Web-UI, CLI, and Python package, powered by OpenAI's Whisper and its variants (faster… | 66 | 1682 | active |
| isair/jarvis Jarvis is a fully local, offline AI voice assistant for your computer that supports natural conversational interaction, unlimited memory, w… | 84 | 1656 | active |
| LokerL/tts-vue TTS-Vue is a cross-platform desktop text-to-speech application built with Electron, Vue, ElementPlus, and Vite that uses Microsoft's Edge T… | 58 | 6101 | maintenance |
| SevaSk/ecoute Ecoute is a live transcription application that captures audio from both the user's microphone and speakers and displays real-time transcri… | 64 | 6047 | maintenance |
| Lynpoint/CyberVerse CyberVerse is an open-source, self-hosted framework for building real-time, voice-first AI agents with optional digital-human video (talkin… | 63 | 1612 | active |
| mkiol/dsnote Speech Note is a Linux desktop and Sailfish OS application for note taking, reading, and translating text using offline Speech to Text, Tex… | 96 | 1607 | active |
| SAGAR-TAMANG/friday-tony-stark-demo F.R.I.D.A.Y. is a Tony Stark-inspired voice AI assistant consisting of a FastMCP server exposing tools (web search, news, system info) and … | 55 | 1595 | active |
| finnvoor/yap yap is a Swift CLI for on-device speech transcription of audio and video files using Apple's Speech.framework on macOS 26. It also supports… | 81 | 1586 | active |
| royshil/obs-localvocal LocalVocal is an OBS Studio plugin that performs real-time, fully local speech recognition and translation using Whisper models running on … | 81 | 1585 | active |
| bootphon/phonemizer A Python library and command-line tool that converts words and texts into phonemes across many languages. It wraps multiple backends (espea… | 85 | 1567 | stable |
| byjlw/video-analyzer A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs … | 60 | 1559 | active |
| Enemyx-net/VibeVoice-ComfyUI A ComfyUI custom node integration for Microsoft's VibeVoice text-to-speech model, providing single and multi-speaker voice synthesis with v… | 53 | 1549 | active |
| lucidrains/soundstorm-pytorch A PyTorch implementation of SoundStorm, Google DeepMind's efficient parallel audio generation model that applies MaskGiT-style masked gener… | 39 | 1546 | active |
| RunanywhereAI/RCLI RCLI is a single-binary CLI that runs open-source AI models locally on your machine, covering chat, vision, speech-to-text, text-to-speech,… | 82 | 1542 | active |
| vibevoice-community/VibeVoice VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker … | 60 | 1542 | active |
| pipecat-ai/smart-turn Smart Turn is an open-source, audio-native turn detection model that decides when a voice agent should respond to human speech, using proso… | 49 | 1538 | active |
| bytedance/SALMONN SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video… | 73 | 1513 | active |
| kyutai-labs/hibiki Hibiki is a decoder-only model for streaming (simultaneous) speech-to-speech translation, built on the multistream Moshi architecture. It p… | 28 | 1509 | active |
| amicalhq/amical Amical is an open-source, local-first AI dictation app for macOS and Windows that transcribes speech with Whisper and enhances it with LLMs… | 84 | 1506 | active |
| kyutai-labs/unmute Unmute is a system that lets any text LLM listen and speak by wrapping it with Kyutai's low-latency speech-to-text and text-to-speech model… | 60 | 1506 | active |
| Flashlight wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into … | 62 | 5466 | maintenance |
| stepfun-ai/Step-Audio2 Step-Audio 2 is an end-to-end multimodal large language model for industry-strength audio understanding and speech conversation, with open-… | 50 | 1503 | active |
| QwenAudio/Fun-ASR Fun-ASR is a family of open-source LLM-based end-to-end speech recognition models from Tongyi Lab, covering Chinese, dialects, accents, and… | 81 | 1496 | active |
| ibab/tensorflow-wavenet A TensorFlow implementation of DeepMind's WaveNet generative neural network architecture for raw audio waveform generation. It provides tra… | 32 | 5428 | maintenance |
| espressif/esp-sr ESP-SR is Espressif's speech recognition framework for ESP32-series chips, providing wake word detection (WakeNet), voice activity detectio… | 78 | 1492 | active |
| ekwek1/soprano Soprano is an ultra-lightweight 80M-parameter text-to-speech model and Python library for fast, expressive, high-fidelity speech synthesis … | 44 | 1486 | active |
| fossasia/voxbento Voxbento is an open-source, self-hosted platform for real-time simultaneous interpretation at live events. Interpreters broadcast translate… | 77 | 1482 | active |
| k2-fsa/icefall Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi… | 64 | 1482 | active |
| IBM Watson Node.js SDK The official Node.js/TypeScript client library (npm package ibm-watson) for accessing IBM Watson services such as Assistant, Speech to Text… | 84 | 1473 | active |
| McCloudS/subgen Subgen is a self-hosted service that automatically generates subtitles for media files using OpenAI's Whisper speech recognition models. It… | 74 | 1465 | active |
| NVIDIA/tacotron2 NVIDIA's PyTorch implementation of the Tacotron 2 text-to-speech model, which synthesizes mel spectrograms from text for vocoder-based audi… | 32 | 5296 | maintenance |
| microsoft/NeuralSpeech NeuralSpeech is a Microsoft Research Asia research repository containing implementations of neural speech processing models across ASR erro… | 32 | 1462 | active |
| silverstein/minutes Minutes is an open-source, privacy-first conversation memory app that records meetings, voice memos, and dictation, transcribes them locall… | 77 | 1453 | active |
| DicioTeam/dicio-android Dicio is a free and open source voice assistant app for Android that interprets user questions and provides speech and graphical feedback e… | 83 | 1452 | active |
| jxlpzqc/TMSpeech TMSpeech is a Windows desktop application that captures system audio via WASAPI loopback and performs real-time Chinese speech-to-text, dis… | 82 | 1447 | active |
| voice-cloning-app/Voice-Cloning-App A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom… | 23 | 1440 | active |
| edwko/OuteTTS OuteTTS is a Python interface for running OuteAI's text-to-speech models, supporting multiple inference backends including llama.cpp, Huggi… | 52 | 1436 | active |
| FireRedTeam/FireRedTTS2 FireRedTTS-2 is a long-form streaming text-to-speech system for multi-speaker dialogue generation, built in PyTorch with a dual-transformer… | 40 | 1428 | active |
| devnen/Chatterbox-TTS-Server A self-hosted server wrapping Resemble AI's Chatterbox TTS models behind an OpenAI-compatible API with a modern web UI. It supports voice c… | 65 | 1420 | active |
| lenML/Speech-AI-Forge Speech-AI-Forge is a Python project built around multiple TTS generation models (ChatTTS, CosyVoice, Fish-Speech, Index-TTS, F5-TTS, FireRe… | 62 | 1416 | active |
| joewongjc/type4me Type4Me is a macOS voice input method that transcribes speech in real time (streaming, local or cloud ASR engines) and refines text with LL… | 81 | 1404 | active |
| 0nutation/SpeechGPT SpeechGPT is a series of speech large language models that can perceive and generate speech, supporting cross-modal instruction following a… | 29 | 1401 | active |
| zackees/transcribe-anything A Python CLI app that transcribes local audio/video files or URLs (YouTube, Rumble, etc.) using multiple Whisper backends with automatic de… | 93 | 1393 | active |
| wenet-e2e/wespeaker WeSpeaker is a research and production-oriented toolkit for speaker embedding learning, supporting speaker verification, recognition, and d… | 64 | 1392 | active |
| haoheliu/voicefixer VoiceFixer is a Python library and CLI tool for general speech restoration, using a pretrained neural vocoder to restore degraded human spe… | 26 | 1373 | stable |
| MoonInTheRiver/DiffSinger Official PyTorch implementation of DiffSinger, an AAAI 2022 paper on singing voice synthesis and text-to-speech using a shallow diffusion m… | 65 | 4851 | maintenance |
| k2-fsa/k2 k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)… | 64 | 1352 | active |
| nyrahealth/CrisperWhisper CrisperWhisper 2.0 is a controllable speech recognition model and Python library that transcribes audio either verbatim (including fillers,… | 89 | 1349 | active |
| shivammehta25/Matcha-TTS Matcha-TTS is a PyTorch-based text-to-speech system that uses conditional flow matching for fast, non-autoregressive speech synthesis. It s… | 62 | 1349 | active |
| Softcatala/whisper-ctranslate2 A command-line transcription and translation tool compatible with OpenAI's Whisper CLI, built on CTranslate2 and faster-whisper for up to 4… | 60 | 1343 | active |
| claritylab/lucida Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros… | 32 | 4782 | maintenance |
| mbailey/voicemode VoiceMode is a Python-based MCP server and Claude Code plugin that enables natural, hands-free voice conversations with Claude Code and oth… | 80 | 1339 | active |
| mmorise/World WORLD is a C++ library for high-quality speech analysis, manipulation, and synthesis based on a vocoder design. It estimates F0 (via DIO/Ha… | 64 | 1338 | stable |
| google-gemma/gemma-translator An on-device, fully offline voice translator application powered by Gemma 4 via LiteRT-LM, with a React web frontend optimized for small ha… | 56 | 1337 | active |
| joey-zhou/xiaozhi-esp32-server-java A Java enterprise-grade server and management platform for the Xiaozhi ESP32 AI voice assistant hardware, providing a full front-end/back-e… | 69 | 1336 | active |
| GauravSingh9356/J.A.R.V.I.S A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,… | 48 | 1332 | active |
| andimarafioti/faster-qwen3-tts A Python library for real-time text-to-speech inference with Qwen3-TTS using manual CUDA graph capture, requiring no Flash Attention, vLLM,… | 78 | 1330 | active |
| sanchit-gandhi/whisper-jax An optimized JAX implementation of OpenAI's Whisper speech recognition model, built on Hugging Face Transformers, offering up to 70x faster… | 30 | 4682 | maintenance |
| jishengpeng/WavTokenizer WavTokenizer is a state-of-the-art discrete neural audio codec that compresses speech, music, and general audio into only 40 or 75 discrete… | 27 | 1316 | active |
| yeyupiaoling/VoiceprintRecognition-Pytorch A PyTorch-based voiceprint recognition (speaker recognition) framework implementing models such as ECAPA-TDNN, ResNetSE, ERes2Net, and CAM+… | 58 | 1312 | active |
| Henry-23/VideoChat A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late… | 48 | 1303 | active |
| stenolabs/stenoai Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire… | 84 | 1299 | active |
| ictnlp/StreamSpeech StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p… | 37 | 1287 | active |