function: speech-recognition
801 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| goodroot/hyprwhspr hyprwhspr is a native Linux system-wide speech-to-text dictation application supporting local models (Whisper, Parakeet, Cohere) with optio… | 85 | 1182 | active |
| NVIDIA/audio-flamingo NVIDIA's PyTorch implementation of the Audio Flamingo series of large audio-language models (AF1, AF2, AF3, and Music Flamingo) for audio u… | 50 | 1182 | active |
| ddlBoJack/emotion2vec Official PyTorch implementation of emotion2vec, a self-supervised pre-trained model for speech emotion representation. It provides code for… | 27 | 1179 | active |
| rudrankriyam/Foundation-Models-Framework-Lab A native iOS and macOS workbench app for learning, testing, and evaluating Apple's Foundation Models framework. It provides editable recipe… | 82 | 1177 | active |
| nari-labs/dia2 Dia2 is a streaming dialogue text-to-speech model from Nari Labs that can generate conversational audio in realtime without needing the ful… | 41 | 1172 | active |
| matthiasn/lotti Lotti is a private, local-first logbook app for journaling, task management, time tracking, habits, and health data, with a staff of person… | 98 | 1169 | active |
| wladradchenko/wunjo.wladradchenko.ru Wunjo CE is an open-source, locally-run AI media suite for face swap, lip sync, voice cloning, object/text/background removal, restyling, a… | 70 | 1169 | active |
| soniqo/speech-swift An open-source Swift toolkit for on-device speech AI on Apple Silicon, providing ASR, TTS, speech-to-speech, VAD, and speaker diarization v… | 83 | 1161 | active |
| GENEXIS-AI/chromex Chromex is a Chrome MV3 side-panel extension that connects the browser to OpenAI's Codex CLI via a local native-messaging bridge. It lets u… | 72 | 1152 | active |
| janvarev/Irene-Voice-Assistant Irene is an offline-capable Russian voice assistant written in Python that recognizes speech via Vosk and responds using TTS engines. Its f… | 65 | 1152 | active |
| TensorSpeech/TensorFlowTTS TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me… | 23 | 3995 | maintenance |
| google/lyra Lyra is a very low-bitrate speech codec from Google that combines traditional codec techniques with generative machine learning models to c… | 23 | 3973 | maintenance |
| midas-research/audino Audino is an open-source, self-hosted web application for annotating audio, supporting transcription, labeling, and speaker-related tasks. … | 52 | 1146 | active |
| facebookresearch/fairseq2 fairseq2 is a PyTorch-based sequence modeling toolkit from Meta FAIR for training custom models for content generation tasks such as langua… | 89 | 1143 | active |
| andabi/deep-voice-conversion A TensorFlow implementation of deep neural networks for voice conversion (voice style transfer) that converts a source speaker's voice into… | 32 | 3938 | maintenance |
| xzf-thu/Mega-ASR Mega-ASR is a foundation automatic speech recognition model trained on 2.6M samples spanning 7 atomic acoustic conditions and 54 compound r… | 59 | 1136 | active |
| video-db/call.md Call.md is an open-source Electron desktop app that records meetings locally, transcribes them in real time with speaker separation, and pr… | 59 | 1132 | active |
| sooftware/conformer An unofficial PyTorch implementation of the Conformer architecture (convolution-augmented Transformer) from the INTERSPEECH 2020 paper, tar… | 72 | 1131 | active |
| PriesiaMioShirakana/DragonianVoice A C++ inference library for running ONNX-based TTS, SVC (singing voice conversion), and SVS (singing voice synthesis) models, supporting ar… | 40 | 1129 | active |
| PromtEngineer/Verbi Verbi is a modular Python voice assistant application for experimenting with state-of-the-art transcription, LLM response generation, and t… | 48 | 1124 | active |
| FutureUniant/Tailor Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea… | 37 | 1115 | active |
| ardha27/AI-Waifu-Vtuber An AI waifu Vtuber application that listens to your voice via Whisper speech recognition, generates in-character replies with OpenAI, and s… | 68 | 1113 | active |
| KoljaB/RealtimeVoiceChat A Python client-server application enabling natural spoken conversations with LLMs, streaming browser audio via WebSockets through Realtime… | 33 | 3832 | maintenance |
| gabber-dev/gabber Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr… | 44 | 1111 | active |
| fxy2311-youyou/expression-trainer A local desktop app (Electron) that trains spoken expression skills by combining fully offline real-time speech recognition (Sherpa-ONNX) w… | 55 | 1105 | active |
| dominostars/playtranslate PlayTranslate is a real-time screen translation app for Android that captures game or app text via OCR and translates it, with support for … | 81 | 1102 | active |
| LoredCast/filewizard A self-hosted, browser-based web UI for converting files between many formats, running OCR on PDFs and images, transcribing audio with Whis… | 46 | 1102 | active |
| vocodedev/vocode-core Vocode is an open-source Python library for building real-time, voice-based LLM applications and agents. It orchestrates streaming transcri… | 21 | 3785 | maintenance |
| Xinrea/bili-shadowreplay A desktop and Docker-deployable tool that continuously caches live streams (Bilibili, Douyin, TikTok, Kuaishou) and lets users clip segment… | 93 | 1099 | active |
| savbell/whisper-writer WhisperWriter is a small desktop dictation app that transcribes microphone audio to text using OpenAI's Whisper model, either locally via f… | 20 | 1097 | active |
| Saik0s/Whisperboard WhisperBoard is an open-source iOS app for voice recording and transcription powered by OpenAI's Whisper model running on-device via whispe… | 57 | 1093 | active |
| nobodywho-ooo/nobodywho NobodyWho is an open-source (EUPL-1.2) on-device LLM inference engine written in Rust, built on llama.cpp, with SDKs for Kotlin, Swift, Pyt… | 86 | 1086 | active |
| FujiwaraChoki/supoclip SupoClip is an open-source, AI-powered video clipping tool that turns long videos like podcasts and streams into vertical 9:16 short clips … | 62 | 1086 | active |
| olivia-ai/olivia Olivia is an open-source chatbot written in Go that uses a neural network for natural language understanding, aiming to be a free alternati… | 10 | 3717 | maintenance |
| dtsola/xiaoyaosearch XiaoyaoSearch is a cross-platform desktop application (Electron + Python/FastAPI) that lets users find local files using AI-powered semanti… | 70 | 1083 | active |
| festvox/flite Flite (festival-lite) is a small, fast, portable run-time text-to-speech synthesis engine written entirely in ANSI C by Carnegie Mellon Uni… | 23 | 1081 | stable |
| XiaomiMiMo/MiMo-Audio Xiaomi's open-source 7B audio language model family (Base and Instruct) plus a 1.2B RVQ audio tokenizer, pretrained on 100M+ hours of audio… | 57 | 1077 | active |
| Muesli-HQ/muesli Muesli is an open-source native macOS app that combines hotkey-driven AI dictation with local meeting transcription, running speech-to-text… | 82 | 1072 | active |
| ufal/whisper_streaming A Python library that turns Whisper-like speech recognition models into a real-time streaming transcription and translation system using a … | 53 | 3672 | maintenance |
| IAHispano/Applio Applio is an open-source, MIT-licensed voice conversion suite built on RVC that lets users convert audio into other voices, train custom vo… | 90 | 3644 | maintenance |
| brenpoly/be-more-agent An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp… | 64 | 1060 | active |
| Makememo/MemoAI MemoAI is a desktop application for macOS and Windows that transcribes audio and video (YouTube links, podcasts, local files) into text and… | 96 | 1059 | active |
| AnnaSuSu/TechSpar TechSpar is an open-source AI-powered technical interview preparation application that combines targeted training, resume-based mock interv… | 82 | 1055 | active |
| zai-org/GLM-TTS GLM-TTS is a Python-based text-to-speech synthesis system built on large language models, using a two-stage LLM plus Flow model architectur… | 50 | 1055 | active |
| alphacep/vosk-android-demo A demo Android application showing offline speech recognition and speaker identification using the Vosk and Kaldi libraries. It serves as a… | 48 | 1055 | active |
| Artrajz/vits-simple-api A Python HTTP API service that exposes VITS-family text-to-speech models (VITS, Bert-VITS2, GPT-SoVITS, W2V2 emotional VITS) for inference.… | 69 | 1052 | active |
| introlab/odas ODAS (Open embeddeD Audition System) is a C library for real-time sound source localization, tracking, separation, and post-filtering using… | 32 | 1048 | stable |
| Kieirra/murmure Murmure is a privacy-first, open-source desktop speech-to-text application that transcribes voice entirely on-device using NVIDIA's Parakee… | 85 | 1047 | active |
| shang-zhu/violin Violin is an open-source video translation tool that transcribes speech, translates it into 33 languages, synthesizes a native-sounding voi… | 52 | 1047 | active |
| xiaochong/hi-kid HiKid is a free, open-source Electron desktop app that lets children in non-English-speaking countries practice English speaking and listen… | 54 | 1043 | active |
| PatterAI/Patter Patter is an open-source, MIT-licensed SDK (Python and TypeScript) that connects AI agents to real phone calls, handling telephony, speech-… | 78 | 1039 | active |
| susiai/susi_chat A chat interface for communicating with a locally hosted LLM (via llama.cpp) through a terminal console, a browser-based console, or a voic… | 61 | 1036 | active |
| mybigday/llama.rn A React Native binding of llama.cpp that enables on-device LLM inference on iOS and Android. It supports GPU/NPU acceleration (Metal, Hexag… | 94 | 1026 | active |
| rayenfeng/riko_project Project Riko is an anime-themed conversational voice assistant that combines OpenAI's GPT for dialogue, GPT-SoVITS for voice synthesis, and… | 31 | 1023 | active |
| hezarai/hezar Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta… | 78 | 1013 | active |
| TensorSpeech/TensorFlowASR TensorFlowASR is a Python library implementing automatic speech recognition architectures such as DeepSpeech2, Jasper, RNN Transducer, Cont… | 66 | 1010 | active |
| HumeAI/tada TADA is an open-source speech-language model from Hume AI that generates expressive, high-fidelity speech via text-acoustic dual alignment,… | 51 | 1009 | active |
| voquill/voquill Voquill is an open-source, cross-platform AI voice dictation app that lets users dictate into any desktop application, with AI-powered tran… | 77 | 1008 | active |
| InsiderX-Pro/video-translator An OpenClaw agent skill (written in Shell/Python) that translates and dubs videos by submitting jobs to a remote video-translation service … | 54 | 1007 | active |
| AutoArk/open-audio-opd An industrial training stack for online policy distillation (OPD) of audio models, distilling compact ASR (and planned TTS) student models … | 52 | 1007 | active |
| octimot/StoryToolkitAI StoryToolkitAI is a desktop film editing tool that transcribes, indexes, and semantically searches video footage locally, using speech reco… | 65 | 1006 | active |
| Jingyi-Wu-Richael/rachel-digital-human-production A Codex skill (plugin) that packages a repeatable workflow for producing authorized digital-human talking-head videos using MiniMax voice c… | 54 | 1005 | active |
| fossasia/voxbento_translator A Flask-based HTTP API for real-time speech-to-text transcription and translation that accepts streamed audio chunks and routes them to plu… | 75 | 1002 | active |
| resemble-ai/Resemblyzer Resemblyzer is a Python package that uses a deep learning voice encoder to convert speech audio into 256-dimensional voice embeddings. Thes… | 23 | 3300 | maintenance |
| carykh/jumpcutter A Python CLI tool that automatically edits videos by speeding up or cutting silent sections using audio volume analysis and ffmpeg. This Gi… | 32 | 3149 | maintenance |
| pluja/whishper Whishper is a self-hosted, 100% local audio transcription and subtitling suite with a web UI, powered by FasterWhisper. It transcribes audi… | 61 | 3066 | maintenance |
| rishikanthc/Scriberr Scriberr is an open-source, fully offline audio transcription application designed for self-hosters who value privacy. It transcribes audio… | 70 | 2990 | maintenance |
| HaujetZhao/QuickCut QuickCut is a lightweight, open-source video processing application built with Python and PyQt that wraps FFmpeg in a friendly GUI. It hand… | 23 | 2949 | maintenance |
| pytorch/audio TorchAudio is PyTorch's audio library providing data manipulation, transforms, and dataset loaders for audio and speech machine learning. I… | 87 | 2927 | maintenance |
| tensorflow/lingvo Lingvo is a TensorFlow-based framework for building neural networks, particularly sequence models, with a focus on speech recognition, mach… | 72 | 2864 | maintenance |
| readbeyond/aeneas aeneas is a Python/C library and set of CLI tools that automatically computes forced alignments, generating a synchronization map between a… | 65 | 2863 | maintenance |
| evancohen/smart-mirror A DIY voice-controlled smart mirror application that displays information and controls IoT smart devices, typically running on a Raspberry … | 32 | 2820 | maintenance |
| marytts/marytts MaryTTS is an open-source, multilingual text-to-speech synthesis platform written in pure Java, operating as a client-server system. It sup… | 24 | 2583 | maintenance |
| zzmp/juliusjs JuliusJS is a JavaScript port of the Julius speech recognition engine that runs entirely in the browser via a Web Worker. It transcribes us… | 23 | 2565 | maintenance |
| s3prl/s3prl S3PRL is a PyTorch toolkit for self-supervised speech pre-training and representation learning, bundling many upstream models like wav2vec … | 55 | 2561 | maintenance |
| wiseman/py-webrtcvad A Python wrapper around Google's WebRTC Voice Activity Detector, classifying short frames of 16-bit mono PCM audio as speech or non-speech.… | 32 | 2496 | maintenance |
| CjangCjengh/MoeGoe MoeGoe is an executable command-line tool for running inference with VITS text-to-speech models, supporting TTS, voice conversion, HuBERT-V… | 23 | 2423 | maintenance |
| jameslyons/python_speech_features A Python library for extracting common speech features used in automatic speech recognition, including MFCCs, filterbank energies, log filt… | 23 | 2423 | maintenance |
| mravanelli/pytorch-kaldi PyTorch-Kaldi is a toolkit for developing state-of-the-art DNN/HMM hybrid speech recognition systems, combining PyTorch-managed neural netw… | 32 | 2404 | maintenance |
| cogentapps/chat-with-gpt An open-source, self-hostable ChatGPT web app built with TypeScript and React that adds voice features via ElevenLabs text-to-speech and Op… | 30 | 2352 | maintenance |
| jianfch/stable-ts A Python library that modifies OpenAI's Whisper to produce more reliable timestamps, adding transcription, forced alignment, and audio inde… | 10 | 2281 | maintenance |
| m1guelpf/auto-subtitle A Python CLI tool that uses OpenAI's Whisper and ffmpeg to automatically generate subtitles for videos and burn them into the output file. … | 32 | 2275 | maintenance |
| jarikomppa/soloud SoLoud is a free, portable C/C++ audio engine for games with a simple 'fire and forget' API for playing sounds. It supports WAV, Ogg Vorbis… | 23 | 2166 | maintenance |
| SeanNaren/deepspeech.pytorch A PyTorch implementation of the DeepSpeech2 speech recognition model, built on PyTorch Lightning, supporting training, testing, and inferen… | 23 | 2136 | maintenance |
| nobody132/masr MASR is an end-to-end Mandarin Chinese automatic speech recognition project built on a gated convolutional neural network (similar to Wav2L… | 23 | 1968 | maintenance |
| QwenLM/Qwen-Audio Official repository for Qwen-Audio, Alibaba Cloud's large audio-language model with pretrained and chat variants. It provides model weights… | 27 | 1945 | maintenance |
| julius-speech/julius Julius is an open-source large vocabulary continuous speech recognition (LVCSR) decoder written in C, based on word N-gram language models … | 35 | 1933 | maintenance |
| astorfi/lip-reading-deeplearning A TensorFlow implementation of coupled 3D convolutional neural networks for cross audio-visual matching recognition, accompanying an IEEE A… | 23 | 1904 | maintenance |
| kalliope-project/kalliope Kalliope is a modular, always-on voice-controlled personal assistant framework written in Python, designed to run on Linux systems includin… | 23 | 1772 | maintenance |
| strob/gentle Gentle is a robust yet lenient forced aligner built on Kaldi that aligns audio speech with a known text transcript. It can be used as a Mac… | 65 | 1705 | maintenance |
| bjoernkarmann/project_alias Project Alias is an open-source Raspberry Pi-based device that acts as a 'parasite' on smart home assistants, letting users train custom wa… | 32 | 1700 | maintenance |
| google/aiyprojects-raspbian Google's Python API libraries, samples, and Raspbian system images for the AIY Projects Voice Kit and Vision Kit on Raspberry Pi. It provid… | 10 | 1663 | maintenance |
| vanshg/MacAssistant MacAssistant is a macOS application that integrates the Google Assistant into the Mac menu bar using the Google Assistant SDK. It is writte… | 23 | 1603 | maintenance |
| google/uis-rnn A Python library implementing the Unbounded Interleaved-State Recurrent Neural Network (UIS-RNN) algorithm for segmenting and clustering se… | 10 | 1588 | maintenance |
| yakGPT/yakGPT YakGPT is a locally running, browser-based ChatGPT UI that connects directly to the OpenAI API with your own key. It adds hands-free voice … | 30 | 1586 | maintenance |
| fossasia/susi_chromebot A Chrome browser extension that provides access to the SUSI.AI assistant without leaving the current tab. It opens via toolbar icon or Alt+… | 10 | 1546 | maintenance |
| beyondcode/writeout.ai A self-hosted Laravel web application that transcribes uploaded audio files using OpenAI's Whisper API and translates the transcripts via t… | 71 | 1525 | maintenance |
| Tameyer41/liftoff Liftoff is a web application that simulates technical mock interviews and provides AI-powered feedback using OpenAI Whisper for speech tran… | 29 | 1523 | maintenance |
| syl22-00/pocketsphinx.js PocketSphinx.js is a speech recognition library that runs entirely in the web browser, built by compiling the PocketSphinx C recognizer to … | 32 | 1508 | maintenance |
| google/live-transcribe-speech-engine Android client libraries from Google's Live Transcribe app for streaming real-time speech recognition via the Google Cloud Speech API. It p… | 10 | 1499 | maintenance |