function: speech-recognition
801 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| microsoft/markitdown MarkItDown is a lightweight Python utility from Microsoft that converts many file formats (PDF, Office documents, images, audio, HTML, EPub… | 83 | 176487 | active |
| openai/whisper OpenAI's Whisper is a general-purpose speech recognition model and Python library built on a Transformer sequence-to-sequence architecture.… | 68 | 107981 | stable |
| RVC-Boss/GPT-SoVITS GPT-SoVITS is a Python-based few-shot voice cloning and text-to-speech system with an integrated WebUI. It supports zero-shot TTS from a 5-… | 68 | 61255 | active |
| microsoft/VibeVoice VibeVoice is Microsoft's open-source frontier voice AI framework combining a next-token diffusion text-to-speech model for expressive, long… | 60 | 53225 | active |
| ggml-org/whisper.cpp A high-performance C/C++ port of OpenAI's Whisper automatic speech recognition model with no external dependencies. It supports CPU, GPU (M… | 99 | 53203 | stable |
| jamiepine/voicebox Voicebox is a free, open-source, local-first AI voice studio desktop app that combines voice cloning, text-to-speech across 7 engines, and … | 75 | 51551 | active |
| moeru-ai/airi Project AIRI is an open-source, self-hosted AI companion (waifu/virtual character) platform inspired by Neuro-sama, supporting Live2D and V… | 86 | 48451 | active |
| 2noise/ChatTTS ChatTTS is a generative text-to-speech model optimized for conversational dialogue scenarios such as LLM assistants, supporting English and… | 69 | 39797 | active |
| RVC-Project/Retrieval-based-Voice-Conversion-WebUI A WebUI-based framework for training and running retrieval-based voice conversion (RVC) models, letting users clone a voice timbre from as … | 82 | 37845 | active |
| myshell-ai/OpenVoice OpenVoice is a Python library and audio foundation model for instant voice cloning, requiring only a short reference audio clip to replicat… | 34 | 37312 | active |
| ZuodaoTech/everyone-can-use-english Everyone Can Use English is an open-source project combining a well-known English learning book with the Enjoy app, an AI-powered language … | 74 | 36774 | active |
| OpenBMB/VoxCPM VoxCPM is a tokenizer-free text-to-speech system built on a diffusion autoregressive architecture that generates continuous speech represen… | 79 | 36145 | active |
| fishaudio/fish-speech Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 l… | 74 | 32413 | active |
| cjpais/Handy Handy is a free, open-source, cross-platform desktop application for offline speech-to-text dictation. Press a configurable shortcut, speak… | 83 | 30402 | active |
| Zackriya-Solutions/meetily Meetily is a privacy-first, open-source AI meeting assistant that records, transcribes, and summarizes meetings entirely locally using Para… | 79 | 29940 | active |
| resemble-ai/chatterbox Chatterbox is a family of state-of-the-art open-source text-to-speech models by Resemble AI, including multilingual and low-latency Turbo v… | 52 | 26163 | active |
| mozilla-ai/llamafile llamafile is a Mozilla project that packages llama.cpp (and whisper.cpp for whisperfile) with Cosmopolitan Libc into single-file executable… | 94 | 25694 | active |
| SYSTRAN/faster-whisper A fast reimplementation of OpenAI's Whisper speech-to-text model built on the CTranslate2 inference engine, offering up to 4x speedup and l… | 57 | 25104 | active |
| m-bain/whisperX WhisperX is a Python library providing fast automatic speech recognition with word-level timestamps and speaker diarization. It combines Wh… | 91 | 23765 | active |
| index-tts/index-tts IndexTTS is an industrial-level zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest IndexTTS… | 78 | 23511 | active |
| QwenAudio/CosyVoice CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-… | 61 | 22925 | active |
| browser-use/video-use An open-source agent skill that lets coding agents like Claude Code edit raw video footage into finished MP4s via natural-language chat. It… | 59 | 21397 | active |
| chidiwilliams/buzz Buzz is a cross-platform desktop application that transcribes and translates audio and video offline using OpenAI's Whisper models. It offe… | 95 | 21147 | active |
| w-okada/voice-changer VCClient is a real-time AI voice changer application that converts microphone input using models like RVC and Beatrice v2. It ships as preb… | 65 | 20864 | active |
| modelscope/FunASR FunASR is an industrial end-to-end speech recognition toolkit built on PyTorch, offering ASR, VAD, punctuation restoration, speaker diariza… | 99 | 20036 | active |
| jianchang512/pyvideotrans pyVideoTrans is an open-source desktop application (with WebUI and CLI modes) that translates videos from one language to another. It provi… | 90 | 18812 | active |
| NVIDIA-NeMo/Speech NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, … | 98 | 18337 | active |
| Huanshere/VideoLingo VideoLingo is an all-in-one AI video translation, localization, and dubbing tool that produces Netflix-quality single-line subtitles with w… | 80 | 18266 | active |
| leon-ai/leon Leon is an open-source personal AI assistant that runs on your own server, combining tools, memory, and agentic execution with speech-to-te… | 86 | 17462 | active |
| bradautomates/claude-video A Claude Code plugin and Agent Skill that lets AI coding agents watch videos. It downloads videos via yt-dlp, extracts scene-aware frames w… | 74 | 16260 | active |
| WEIFENG2333/VideoCaptioner VideoCaptioner is an open-source AI-powered video subtitling tool that handles the full pipeline: speech recognition, LLM-based subtitle se… | 82 | 15764 | active |
| KittenML/KittenTTS KittenTTS is an open-source, ultra-lightweight text-to-speech library built on ONNX, with models ranging from 15M to 80M parameters (25-80 … | 66 | 15403 | active |
| Vosk Vosk is an offline open source speech recognition toolkit supporting 20+ languages with small models and a streaming API. It provides bindi… | 66 | 15077 | stable |
| duixcom/Duix-Avatar Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea… | 57 | 14871 | active |
| SesameAILabs/csm CSM (Conversational Speech Model) is Sesame's speech generation model that produces conversational audio from text and audio context, using… | 30 | 14720 | active |
| k2-fsa/sherpa-onnx sherpa-onnx is an offline speech processing toolkit built on next-gen Kaldi and onnxruntime, supporting speech-to-text, text-to-speech, spe… | 95 | 14411 | active |
| SubtitleEdit/subtitleedit Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac… | 94 | 13971 | active |
| Open-LLM-VTuber/Open-LLM-VTuber Open-LLM-VTuber is a cross-platform AI companion application that enables real-time voice conversations with LLMs, featuring a Live2D avata… | 67 | 13473 | active |
| Omi Omi is an open-source AI wearable and companion app ecosystem (necklace pendant, smart glasses, desktop and mobile apps) that captures conv… | 88 | 13264 | active |
| livekit/agents LiveKit Agents is an open-source framework for building realtime, multimodal voice AI agents that run as programmable participants in LiveK… | 90 | 13176 | active |
| QwenLM/Qwen3-TTS Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,… | 48 | 13113 | active |
| Vaibhavs10/insanely-fast-whisper A CLI tool for transcribing audio files on-device using OpenAI's Whisper models, powered by Hugging Face Transformers, Optimum, and Flash A… | 49 | 13049 | active |
| huggingface/speech-to-speech A low-latency, fully modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) built from open-source models, exposed via the OpenAI Realtime… | 85 | 12882 | active |
| PaddlePaddle/PaddleSpeech PaddleSpeech is an open-source speech and audio toolkit built on the PaddlePaddle deep learning platform, covering ASR with punctuation, st… | 66 | 12670 | active |
| abus-aikorea/voice-pro Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, … | 84 | 12647 | active |
| facebookresearch/seamless_communication A library of foundational multilingual multimodal AI models from Meta for speech and text translation, including SeamlessM4T, SeamlessExpre… | 71 | 11842 | active |
| rany2/edge-tts A Python module and CLI that lets you use Microsoft Edge's online neural text-to-speech service without needing Edge, Windows, or an API ke… | 79 | 11803 | active |
| speechbrain/speechbrain SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained… | 83 | 11785 | active |
| debpalash/VoiceStudio VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, vid… | 80 | 11737 | active |
| ankidroid/Anki-Android AnkiDroid is the Android client for the open-source Anki spaced repetition flashcard system, letting users study and sync decks with AnkiWe… | 93 | 11628 | active |
| CorentinJ/Real-Time-Voice-Cloning A Python implementation of the SV2TTS (Transfer Learning from Speaker Verification to Multispeaker TTS) framework that clones a voice from … | 64 | 60110 | maintenance |
| krillinai/KrillinAI KrillinAI is an AI-powered video translation and dubbing tool covering the full pipeline: video download, speech transcription, subtitle tr… | 84 | 11272 | active |
| TEN-framework/ten-framework TEN is an open-source framework for building real-time multimodal conversational AI agents, with a focus on low-latency voice assistants. I… | 85 | 11082 | active |
| SparkAudio/Spark-TTS Spark-TTS is an LLM-based text-to-speech system built on Qwen2.5 that generates speech directly from single-stream decoupled speech tokens.… | 27 | 11006 | active |
| QuentinFuxa/WhisperLiveKit WhisperLiveKit is a self-hosted, ultra-low-latency real-time speech-to-text pipeline built on state-of-the-art simultaneous speech research… | 87 | 10962 | active |
| kyutai-labs/moshi Moshi is a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai, built around the Mimi streaming neural audio… | 62 | 10949 | active |
| altic-dev/FluidVoice FluidVoice is an open-source (GPLv3) macOS dictation app that transcribes speech to text entirely on-device using models like Parakeet, Whi… | 84 | 10946 | active |
| moonshine-ai/moonshine Moonshine Voice is an open-source on-device AI toolkit providing very low latency speech-to-text, intent recognition, and text-to-speech fo… | 89 | 10944 | active |
| pyannote/pyannote-audio pyannote.audio is an open-source Python toolkit built on PyTorch for speaker diarization, providing neural building blocks like voice activ… | 95 | 10475 | active |
| xinnan-tech/xiaozhi-esp32-server A self-hosted backend server for the xiaozhi-esp32 open-source smart hardware project, implementing the Xiaozhi communication protocol over… | 81 | 10433 | active |
| NVIDIA/personaplex PersonaPlex is a real-time, full-duplex speech-to-speech conversational model from NVIDIA that supports persona control via text role promp… | 47 | 10393 | active |
| nexmoe/VidBee VidBee is a free, open-source Electron desktop app for downloading video and audio from 1000+ sites (via yt-dlp) and importing local media.… | 84 | 10386 | active |
| niedev/RTranslator RTranslator is a free, open-source, offline real-time translation app for Android that runs speech recognition (Whisper) and translation (M… | 85 | 10355 | active |
| RunanywhereAI/runanywhere-sdks RunAnywhere is a set of cross-platform SDKs (Swift, Kotlin, React Native, Flutter, TypeScript, C++) over a shared C++ core for running AI m… | 85 | 10282 | active |
| KoljaB/RealtimeSTT RealtimeSTT is a Python library for low-latency, real-time speech-to-text with voice activity detection, wake word activation, and streamin… | 94 | 10080 | active |
| snakers4/silero-vad Silero VAD is a pre-trained, enterprise-grade Voice Activity Detector model available via PyPI, runnable with PyTorch or ONNX Runtime. It d… | 86 | 10061 | stable |
| espnet/espnet ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance… | 88 | 9941 | active |
| xorbitsai/inference Xinference is an open-source model serving platform for deploying LLMs, embedding, speech, image, and multimodal models via a unified OpenA… | 91 | 9523 | active |
| k2-fsa/OmniVoice OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style a… | 79 | 9455 | active |
| dessant/buster Buster is a browser extension that solves reCAPTCHA challenges by completing their audio challenges using speech recognition. It works on C… | 91 | 9276 | active |
| fastrepl/anarlog Anarlog is an open-source, local-first AI meeting notepad and Granola alternative that records device audio without a bot joining the call,… | 84 | 9181 | active |
| QwenAudio/SenseVoice SenseVoice is an open-source speech foundation model (SenseVoiceSmall) providing multilingual ASR, spoken language identification, speech e… | 88 | 9151 | active |
| zyronon/TypeWords TypeWords is a free, open-source web application for learning English vocabulary and memorizing articles through typing practice. It uses t… | 80 | 9007 | active |
| jianchang512/clone-voice A voice cloning tool with a web interface built on the coqui.ai xtts_v2 model, letting users synthesize speech in any voice from text or co… | 10 | 8989 | active |
| Uberi/speech_recognition A Python library for performing speech recognition with support for multiple engines and APIs, both online (Google, Azure, Wit.ai, OpenAI W… | 94 | 8987 | active |
| Foliate Foliate is a free, open-source e-book reader application for Linux built with GTK4 and GJS, supporting EPUB, Mobipocket, Kindle, FB2, CBZ, … | 64 | 8667 | active |
| jasonppy/VoiceCraft VoiceCraft is a token infilling neural codec language model for zero-shot speech editing and text-to-speech on in-the-wild data like audiob… | 63 | 8574 | active |
| hexgrad/kokoro An inference library for Kokoro-82M, an open-weight text-to-speech model with 82 million parameters that delivers quality comparable to lar… | 36 | 8570 | active |
| netease-youdao/EmotiVoice EmotiVoice is an open-source text-to-speech engine supporting English and Chinese with over 2000 voices and prompt-controlled emotional syn… | 17 | 8523 | active |
| nl8590687/ASRT_SpeechRecognition ASRT is a deep-learning-based Chinese speech recognition (speech-to-text) system built with TensorFlow/Keras, using CNN, LSTM, attention me… | 57 | 8383 | active |
| duixcom/Duix-Mobile Duix Mobile is an open-source SDK for building real-time interactive AI avatars (digital humans) that run on-device on Android, iOS, tablet… | 69 | 8198 | active |
| GetStream/Vision-Agents An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O… | 84 | 8100 | active |
| fonoster/fonoster Fonoster is an open-source programmable telecommunications stack (a Twilio alternative) for building voice and messaging applications, with… | 99 | 8080 | active |
| smacke/ffsubsync FFsubsync is a Python command-line tool that automatically synchronizes subtitle files (e.g., SRT) with video or audio, using voice activit… | 84 | 7854 | active |
| Blaizzy/mlx-audio MLX-Audio is a Python library built on Apple's MLX framework for fast text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS… | 88 | 7793 | active |
| babysor/MockingBird MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real… | 54 | 36909 | maintenance |
| myshell-ai/MeloTTS MeloTTS is a high-quality multi-lingual text-to-speech library supporting English (multiple accents), Spanish, French, Chinese, Japanese, a… | 16 | 7608 | active |
| pickle-com/glass Glass by Pickle is an open-source Electron desktop app that acts as a 'digital mind extension', seeing your screen and listening in real ti… | 42 | 7589 | active |
| farzaa/clicky Clicky is an open-source macOS app that acts as an AI teacher living next to your cursor — it can see your screen, talk with you via voice,… | 50 | 7394 | active |
| ilyhalight/voice-over-translation A userscript/browser extension that adds AI voice-over translation and subtitles to videos on many websites, powered by Yandex services. It… | 98 | 7335 | active |
| Zyphra/Zonos Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I… | 24 | 7244 | active |
| thewh1teagle/vibe Vibe is a cross-platform desktop app for fully offline audio and video transcription using Whisper, Nemotron, and Parakeet models. It suppo… | 93 | 7206 | active |
| JefferyHcool/BiliNote BiliNote is an open-source AI video note-taking application that turns video links from Bilibili, YouTube, Douyin, and local files into str… | 86 | 7169 | active |
| wzpan/wukong-robot wukong-robot is a modular Chinese-language voice assistant / smart speaker project in Python that combines offline wake-word detection, ASR… | 32 | 7126 | active |
| ddean2009/MoneyPrinterPlus A Python desktop application that uses AI LLMs to batch-generate short videos with one click, including automatic video mashup/remixing and… | 30 | 7008 | active |
| openai/openai-realtime-agents A Next.js TypeScript demo application showcasing advanced agentic patterns (Chat-Supervisor and Sequential Handoff) built on the OpenAI Rea… | 48 | 6966 | active |
| yihong0618/xiaogpt A Python application that connects Xiaomi AI Speakers to ChatGPT and other LLMs, letting users ask questions starting with a trigger phrase… | 64 | 6901 | active |
| TalAter/annyang annyang is a tiny (2 KB), dependency-free JavaScript library that adds speech recognition and voice commands to websites using the Web Spee… | 78 | 6816 | stable |
| werman/noise-suppression-for-voice A real-time voice noise suppression audio plugin based on Xiph's RNNoise, available in VST2, VST3, LV2, LADSPA, AU, and AUv3 formats. It su… | 77 | 6806 | active |
| Dooy/chatgpt-web-midjourney-proxy A unified web/desktop UI for ChatGPT plus AI image, music, and video generation services like Midjourney, Suno, Luma, Runway, and Flux. It … | 76 | 6785 | active |
page 1 / 9 next →