domain: speech-processing
552 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| YannickJadoul/Parselmouth Parselmouth is a Python library that exposes Praat's internal C/C++ speech and acoustic analysis code through a complete, Pythonic interfac… | 74 | 1286 | stable |
| jasperproject/jasper-client Client code for the Jasper voice computing platform, an open source platform for building always-on, voice-controlled applications. It prov… | 32 | 4518 | maintenance |
| studio-dots-ai/dots.tts dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive text-to-speech system, distributed as a Python library with pretrain… | 79 | 1275 | active |
| phuc-nt/my-translator A Tauri-based desktop app that captures system or microphone audio, transcribes it, and shows translations in a minimal real-time overlay, … | 76 | 1275 | active |
| ABexit/ASR-LLM-TTS An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) … | 60 | 1271 | active |
| yzfly/douyin-mcp-server A Python MCP server and WebUI that extracts watermark-free Douyin (TikTok China) video download links and transcribes video speech into tex… | 10 | 1258 | active |
| peteonrails/voxtype Voxtype is a local, offline voice-to-text dictation tool for Linux with push-to-talk hotkey support across Wayland compositors like Hyprlan… | 78 | 1252 | active |
| gitmylo/audio-webui An all-in-one web UI for audio-related neural networks, bundling text-to-speech (Bark), voice conversion/cloning (RVC), and text-to-audio/m… | 30 | 1247 | active |
| sh-lee-prml/HierSpeechpp Official PyTorch implementation of HierSpeech++, a fast zero-shot speech synthesizer for text-to-speech and voice conversion based on hiera… | 28 | 1238 | active |
| Ksuriuri/index-tts-vllm A reimplementation of IndexTTS's GPT model inference using vLLM, providing significantly faster text-to-speech generation with a web UI and… | 58 | 1235 | active |
| NVIDIA/BigVGAN BigVGAN is NVIDIA's official PyTorch implementation of a universal neural vocoder (ICLR 2023) that generates high-fidelity raw audio wavefo… | 23 | 1227 | stable |
| yeyupiaoling/Whisper-Finetune A toolkit for fine-tuning OpenAI's Whisper speech recognition models using LoRA, supporting training with or without timestamps and even wi… | 66 | 1223 | active |
| Aratako/Irodori-TTS Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer… | 58 | 1219 | active |
| modal-labs/quillman QuiLLMan is a voice chat application built on Kyutai's Moshi speech-to-speech language model, deployed serverlessly on Modal with a FastAPI… | 68 | 1213 | active |
| hgneng/ekho Ekho is an open-source Chinese text-to-speech engine supporting Mandarin, Cantonese, and Tibetan, part of the eGuideDog accessibility proje… | 68 | 1211 | active |
| aTrainTranscription/aTrain aTrain is a desktop GUI application for offline transcription of speech recordings using Whisper-based machine learning models, with speake… | 79 | 1205 | active |
| metavoiceio/metavoice-src MetaVoice-1B is a 1.2B parameter foundational text-to-speech model trained on 100K hours of speech, focused on emotional rhythm and tone in… | 26 | 4205 | maintenance |
| hkjarral/AVA-AI-Voice-Agent-for-Asterisk An open-source AI voice agent that integrates with Asterisk/FreePBX phone systems via Audiosocket/RTP, built in Python with a modular pipel… | 85 | 1202 | active |
| nishuzumi/gemini-teacher A Python CLI application that acts as an English speaking practice assistant powered by Google Gemini. It listens to your speech via microp… | 68 | 1200 | active |
| kizuna-ai-lab/sokuji Sokuji is a cross-platform real-time two-way speech translation app for bilingual meetings, available as a desktop application (Windows, ma… | 83 | 1199 | active |
| diodiogod/TTS-Audio-Suite A ComfyUI custom node suite providing unified multi-engine Text-to-Speech, Voice Conversion, and audio editing across 19 engines like Chatt… | 84 | 1185 | active |
| digimata/parrot Parrot is an ultra-minimalist push-to-talk dictation daemon for macOS that transcribes speech entirely on-device using WhisperKit on the Ap… | 73 | 1185 | active |
| goodroot/hyprwhspr hyprwhspr is a native Linux system-wide speech-to-text dictation application supporting local models (Whisper, Parakeet, Cohere) with optio… | 85 | 1182 | active |
| NVIDIA/audio-flamingo NVIDIA's PyTorch implementation of the Audio Flamingo series of large audio-language models (AF1, AF2, AF3, and Music Flamingo) for audio u… | 50 | 1182 | active |
| ddlBoJack/emotion2vec Official PyTorch implementation of emotion2vec, a self-supervised pre-trained model for speech emotion representation. It provides code for… | 27 | 1179 | active |
| nari-labs/dia2 Dia2 is a streaming dialogue text-to-speech model from Nari Labs that can generate conversational audio in realtime without needing the ful… | 41 | 1172 | active |
| baidu-research/warp-ctc A fast parallel implementation of the Connectionist Temporal Classification (CTC) loss function for CPU and CUDA GPU, with a simple C inter… | 32 | 4069 | maintenance |
| iver56/torch-audiomentations A PyTorch library for fast audio data augmentation, inspired by audiomentations. It provides GPU-accelerated, differentiable audio transfor… | 58 | 1167 | active |
| soniqo/speech-swift An open-source Swift toolkit for on-device speech AI on Apple Silicon, providing ASR, TTS, speech-to-speech, VAD, and speaker diarization v… | 83 | 1161 | active |
| gemelo-ai/vocos Vocos is a fast neural vocoder that synthesizes audio waveforms from acoustic features such as mel-spectrograms or EnCodec tokens. It uses … | 65 | 1155 | stable |
| janvarev/Irene-Voice-Assistant Irene is an offline-capable Russian voice assistant written in Python that recognizes speech via Vosk and responds using TTS engines. Its f… | 65 | 1152 | active |
| TensorSpeech/TensorFlowTTS TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me… | 23 | 3995 | maintenance |
| lhotse-speech/lhotse Lhotse is a Python library for flexible, scalable preparation of multimodal (speech, audio, video, image, text) data for machine learning, … | 89 | 1149 | active |
| google/lyra Lyra is a very low-bitrate speech codec from Google that combines traditional codec techniques with generative machine learning models to c… | 23 | 3973 | maintenance |
| midas-research/audino Audino is an open-source, self-hosted web application for annotating audio, supporting transcription, labeling, and speaker-related tasks. … | 52 | 1146 | active |
| facebookresearch/fairseq2 fairseq2 is a PyTorch-based sequence modeling toolkit from Meta FAIR for training custom models for content generation tasks such as langua… | 89 | 1143 | active |
| andabi/deep-voice-conversion A TensorFlow implementation of deep neural networks for voice conversion (voice style transfer) that converts a source speaker's voice into… | 32 | 3938 | maintenance |
| xzf-thu/Mega-ASR Mega-ASR is a foundation automatic speech recognition model trained on 2.6M samples spanning 7 atomic acoustic conditions and 54 compound r… | 59 | 1136 | active |
| video-db/call.md Call.md is an open-source Electron desktop app that records meetings locally, transcribes them in real time with speaker separation, and pr… | 59 | 1132 | active |
| sooftware/conformer An unofficial PyTorch implementation of the Conformer architecture (convolution-augmented Transformer) from the INTERSPEECH 2020 paper, tar… | 72 | 1131 | active |
| PriesiaMioShirakana/DragonianVoice A C++ inference library for running ONNX-based TTS, SVC (singing voice conversion), and SVS (singing voice synthesis) models, supporting ar… | 40 | 1129 | active |
| PromtEngineer/Verbi Verbi is a modular Python voice assistant application for experimenting with state-of-the-art transcription, LLM response generation, and t… | 48 | 1124 | active |
| ardha27/AI-Waifu-Vtuber An AI waifu Vtuber application that listens to your voice via Whisper speech recognition, generates in-character replies with OpenAI, and s… | 68 | 1113 | active |
| KoljaB/RealtimeVoiceChat A Python client-server application enabling natural spoken conversations with LLMs, streaming browser audio via WebSockets through Realtime… | 33 | 3832 | maintenance |
| gabber-dev/gabber Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr… | 44 | 1111 | active |
| vocodedev/vocode-core Vocode is an open-source Python library for building real-time, voice-based LLM applications and agents. It orchestrates streaming transcri… | 21 | 3785 | maintenance |
| savbell/whisper-writer WhisperWriter is a small desktop dictation app that transcribes microphone audio to text using OpenAI's Whisper model, either locally via f… | 20 | 1097 | active |
| Saik0s/Whisperboard WhisperBoard is an open-source iOS app for voice recording and transcription powered by OpenAI's Whisper model running on-device via whispe… | 57 | 1093 | active |
| festvox/flite Flite (festival-lite) is a small, fast, portable run-time text-to-speech synthesis engine written entirely in ANSI C by Carnegie Mellon Uni… | 23 | 1081 | stable |
| XiaomiMiMo/MiMo-Audio Xiaomi's open-source 7B audio language model family (Base and Instruct) plus a 1.2B RVQ audio tokenizer, pretrained on 100M+ hours of audio… | 57 | 1077 | active |
| Muesli-HQ/muesli Muesli is an open-source native macOS app that combines hotkey-driven AI dictation with local meeting transcription, running speech-to-text… | 82 | 1072 | active |
| ufal/whisper_streaming A Python library that turns Whisper-like speech recognition models into a real-time streaming transcription and translation system using a … | 53 | 3672 | maintenance |
| brenpoly/be-more-agent An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp… | 64 | 1060 | active |
| Makememo/MemoAI MemoAI is a desktop application for macOS and Windows that transcribes audio and video (YouTube links, podcasts, local files) into text and… | 96 | 1059 | active |
| X-LANCE/SLAM-LLM SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc… | 55 | 1056 | active |
| zai-org/GLM-TTS GLM-TTS is a Python-based text-to-speech synthesis system built on large language models, using a two-stage LLM plus Flow model architectur… | 50 | 1055 | active |
| alphacep/vosk-android-demo A demo Android application showing offline speech recognition and speaker identification using the Vosk and Kaldi libraries. It serves as a… | 48 | 1055 | active |
| wxxxcxx/ms-ra-forwarder A self-hostable free online text-to-speech API that proxies Microsoft Edge 'Read Aloud' and Azure TTS demo endpoints. It can be deployed vi… | 61 | 1053 | active |
| Artrajz/vits-simple-api A Python HTTP API service that exposes VITS-family text-to-speech models (VITS, Bert-VITS2, GPT-SoVITS, W2V2 emotional VITS) for inference.… | 69 | 1052 | active |
| introlab/odas ODAS (Open embeddeD Audition System) is a C library for real-time sound source localization, tracking, separation, and post-filtering using… | 32 | 1048 | stable |
| Kieirra/murmure Murmure is a privacy-first, open-source desktop speech-to-text application that transcribes voice entirely on-device using NVIDIA's Parakee… | 85 | 1047 | active |
| shang-zhu/violin Violin is an open-source video translation tool that transcribes speech, translates it into 33 languages, synthesizes a native-sounding voi… | 52 | 1047 | active |
| k2-fsa/ZipVoice ZipVoice is a series of fast, high-quality zero-shot text-to-speech models based on flow matching, with a compact 123M-parameter Zipformer-… | 43 | 1045 | active |
| xiaochong/hi-kid HiKid is a free, open-source Electron desktop app that lets children in non-English-speaking countries practice English speaking and listen… | 54 | 1043 | active |
| PatterAI/Patter Patter is an open-source, MIT-licensed SDK (Python and TypeScript) that connects AI agents to real phone calls, handling telephony, speech-… | 78 | 1039 | active |
| Goekdeniz-Guelmez/Local-NotebookLM A local, open-source alternative to Google's NotebookLM that converts PDF documents into audio content like podcasts, summaries, and interv… | 74 | 1027 | active |
| rayenfeng/riko_project Project Riko is an anime-themed conversational voice assistant that combines OpenAI's GPT for dialogue, GPT-SoVITS for voice synthesis, and… | 31 | 1023 | active |
| hezarai/hezar Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta… | 78 | 1013 | active |
| TensorSpeech/TensorFlowASR TensorFlowASR is a Python library implementing automatic speech recognition architectures such as DeepSpeech2, Jasper, RNN Transducer, Cont… | 66 | 1010 | active |
| HumeAI/tada TADA is an open-source speech-language model from Hume AI that generates expressive, high-fidelity speech via text-acoustic dual alignment,… | 51 | 1009 | active |
| voquill/voquill Voquill is an open-source, cross-platform AI voice dictation app that lets users dictate into any desktop application, with AI-powered tran… | 77 | 1008 | active |
| AutoArk/open-audio-opd An industrial training stack for online policy distillation (OPD) of audio models, distilling compact ASR (and planned TTS) student models … | 52 | 1007 | active |
| fossasia/voxbento_translator A Flask-based HTTP API for real-time speech-to-text transcription and translation that accepts streamed audio chunks and routes them to plu… | 75 | 1002 | active |
| resemble-ai/Resemblyzer Resemblyzer is a Python package that uses a deep learning voice encoder to convert speech audio into 256-dimensional voice embeddings. Thes… | 23 | 3300 | maintenance |
| linyiLYi/bilibot A local chatbot fine-tuned from Bilibili user comments, built on Qwen1.5-32B-Chat using Apple's MLX LoRA fine-tuning. It supports text chat… | 24 | 3139 | maintenance |
| pluja/whishper Whishper is a self-hosted, 100% local audio transcription and subtitling suite with a web UI, powered by FasterWhisper. It transcribes audi… | 61 | 3066 | maintenance |
| keithito/tacotron An unofficial open-source TensorFlow implementation of Google's Tacotron end-to-end neural text-to-speech model, with a pre-trained model a… | 23 | 2995 | maintenance |
| rishikanthc/Scriberr Scriberr is an open-source, fully offline audio transcription application designed for self-hosters who value privacy. It transcribes audio… | 70 | 2990 | maintenance |
| pytorch/audio TorchAudio is PyTorch's audio library providing data manipulation, transforms, and dataset loaders for audio and speech machine learning. I… | 87 | 2927 | maintenance |
| tensorflow/lingvo Lingvo is a TensorFlow-based framework for building neural networks, particularly sequence models, with a focus on speech recognition, mach… | 72 | 2864 | maintenance |
| readbeyond/aeneas aeneas is a Python/C library and set of CLI tools that automatically computes forced alignments, generating a synchronization map between a… | 65 | 2863 | maintenance |
| marytts/marytts MaryTTS is an open-source, multilingual text-to-speech synthesis platform written in pure Java, operating as a client-server system. It sup… | 24 | 2583 | maintenance |
| zzmp/juliusjs JuliusJS is a JavaScript port of the Julius speech recognition engine that runs entirely in the browser via a Web Worker. It transcribes us… | 23 | 2565 | maintenance |
| s3prl/s3prl S3PRL is a PyTorch toolkit for self-supervised speech pre-training and representation learning, bundling many upstream models like wav2vec … | 55 | 2561 | maintenance |
| wiseman/py-webrtcvad A Python wrapper around Google's WebRTC Voice Activity Detector, classifying short frames of 16-bit mono PCM audio as speech or non-speech.… | 32 | 2496 | maintenance |
| jameslyons/python_speech_features A Python library for extracting common speech features used in automatic speech recognition, including MFCCs, filterbank energies, log filt… | 23 | 2423 | maintenance |
| CjangCjengh/MoeGoe MoeGoe is an executable command-line tool for running inference with VITS text-to-speech models, supporting TTS, voice conversion, HuBERT-V… | 23 | 2423 | maintenance |
| mravanelli/pytorch-kaldi PyTorch-Kaldi is a toolkit for developing state-of-the-art DNN/HMM hybrid speech recognition systems, combining PyTorch-managed neural netw… | 32 | 2404 | maintenance |
| r9y9/wavenet_vocoder A PyTorch implementation of the WaveNet vocoder that generates high-quality raw speech waveforms conditioned on acoustic features like mel-… | 23 | 2376 | maintenance |
| NVIDIA/waveglow WaveGlow is a PyTorch implementation of a flow-based generative network that synthesizes high-quality speech audio from mel-spectrograms, c… | 32 | 2339 | maintenance |
| Rayhane-mamah/Tacotron-2 A TensorFlow implementation of DeepMind's Tacotron-2 neural text-to-speech architecture, including both the Tacotron spectrogram predictor … | 32 | 2322 | maintenance |
| jianfch/stable-ts A Python library that modifies OpenAI's Whisper to produce more reliable timestamps, adding transcription, forced alignment, and audio inde… | 10 | 2281 | maintenance |
| fatchord/WaveRNN A PyTorch implementation of DeepMind's WaveRNN neural vocoder plus a Tacotron text-to-speech system, trained on LJSpeech. It supports train… | 32 | 2190 | maintenance |
| ming024/FastSpeech2 A PyTorch implementation of Microsoft's FastSpeech 2 text-to-speech model, supporting English and Mandarin with single- and multi-speaker s… | 32 | 2185 | maintenance |
| SeanNaren/deepspeech.pytorch A PyTorch implementation of the DeepSpeech2 speech recognition model, built on PyTorch Lightning, supporting training, testing, and inferen… | 23 | 2136 | maintenance |
| r9y9/deepvoice3_pytorch A PyTorch implementation of Deep Voice 3 and related convolutional neural network-based text-to-speech synthesis models. It includes traini… | 23 | 1975 | maintenance |
| nobody132/masr MASR is an end-to-end Mandarin Chinese automatic speech recognition project built on a gated convolutional neural network (similar to Wav2L… | 23 | 1968 | maintenance |
| QwenLM/Qwen-Audio Official repository for Qwen-Audio, Alibaba Cloud's large audio-language model with pretrained and chat variants. It provides model weights… | 27 | 1945 | maintenance |
| julius-speech/julius Julius is an open-source large vocabulary continuous speech recognition (LVCSR) decoder written in C, based on word N-gram language models … | 35 | 1933 | maintenance |
| astorfi/lip-reading-deeplearning A TensorFlow implementation of coupled 3D convolutional neural networks for cross audio-visual matching recognition, accompanying an IEEE A… | 23 | 1904 | maintenance |