function: speech-recognition
801 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Lynpoint/CyberVerse CyberVerse is an open-source, self-hosted framework for building real-time, voice-first AI agents with optional digital-human video (talkin… | 63 | 1612 | active |
| mkiol/dsnote Speech Note is a Linux desktop and Sailfish OS application for note taking, reading, and translating text using offline Speech to Text, Tex… | 96 | 1607 | active |
| modem-works/dream-recorder Dream Recorder is an open-source DIY bedside device built on a Raspberry Pi 5 that records users speaking their dreams aloud and turns them… | 41 | 1607 | active |
| Roy3838/Observer Observer AI is a desktop application for building micro-agents that observe screen, camera, microphone, and audio inputs, process them with… | 87 | 1601 | active |
| Norsico/Video-Materials-AutoGEN-Workstation A self-hosted short-video production workstation that combines AI script generation (Gemini), batch TTS voiceover, AI image asset synthesis… | 54 | 1599 | active |
| SAGAR-TAMANG/friday-tony-stark-demo F.R.I.D.A.Y. is a Tony Stark-inspired voice AI assistant consisting of a FastMCP server exposing tools (web search, news, system info) and … | 55 | 1595 | active |
| semperai/amica Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. … | 32 | 1593 | active |
| finnvoor/yap yap is a Swift CLI for on-device speech transcription of audio and video files using Apple's Speech.framework on macOS 26. It also supports… | 81 | 1586 | active |
| royshil/obs-localvocal LocalVocal is an OBS Studio plugin that performs real-time, fully local speech recognition and translation using Whisper models running on … | 81 | 1585 | active |
| TOM88812/xiaozhi-android-client A Flutter-based cross-platform voice chat client for the xiaozhi AI assistant ecosystem, supporting real-time voice interaction and text co… | 60 | 1574 | active |
| bootphon/phonemizer A Python library and command-line tool that converts words and texts into phonemes across many languages. It wraps multiple backends (espea… | 85 | 1567 | stable |
| byjlw/video-analyzer A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs … | 60 | 1559 | active |
| RunanywhereAI/RCLI RCLI is a single-binary CLI that runs open-source AI models locally on your machine, covering chat, vision, speech-to-text, text-to-speech,… | 82 | 1542 | active |
| vibevoice-community/VibeVoice VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker … | 60 | 1542 | active |
| dindin0497/SeeIt SeeIt is an inclusive Android app with two accessibility modes: one that converts spoken speech into text and plays corresponding ASL (Amer… | 41 | 1540 | active |
| tencent-connect/openclaw-qqbot A message channel plugin for OpenClaw that connects an AI assistant to QQ (Tencent's messaging platform), supporting private chat, group ch… | 82 | 1538 | active |
| pipecat-ai/smart-turn Smart Turn is an open-source, audio-native turn detection model that decides when a voice agent should respond to human speech, using proso… | 49 | 1538 | active |
| AlekPet/ComfyUI_Custom_Nodes_AlekPet A collection of custom nodes for ComfyUI that extend its capabilities with painting, pose control, prompt translation, and speech recogniti… | 72 | 1524 | active |
| bytedance/SALMONN SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video… | 73 | 1513 | active |
| kyutai-labs/hibiki Hibiki is a decoder-only model for streaming (simultaneous) speech-to-speech translation, built on the multistream Moshi architecture. It p… | 28 | 1509 | active |
| amicalhq/amical Amical is an open-source, local-first AI dictation app for macOS and Windows that transcribes speech with Whisper and enhances it with LLMs… | 84 | 1506 | active |
| kyutai-labs/unmute Unmute is a system that lets any text LLM listen and speak by wrapping it with Kyutai's low-latency speech-to-text and text-to-speech model… | 60 | 1506 | active |
| Flashlight wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into … | 62 | 5466 | maintenance |
| stepfun-ai/Step-Audio2 Step-Audio 2 is an end-to-end multimodal large language model for industry-strength audio understanding and speech conversation, with open-… | 50 | 1503 | active |
| QwenAudio/Fun-ASR Fun-ASR is a family of open-source LLM-based end-to-end speech recognition models from Tongyi Lab, covering Chinese, dialects, accents, and… | 81 | 1496 | active |
| espressif/esp-sr ESP-SR is Espressif's speech recognition framework for ESP32-series chips, providing wake word detection (WakeNet), voice activity detectio… | 78 | 1492 | active |
| ekwek1/soprano Soprano is an ultra-lightweight 80M-parameter text-to-speech model and Python library for fast, expressive, high-fidelity speech synthesis … | 44 | 1486 | active |
| fossasia/voxbento Voxbento is an open-source, self-hosted platform for real-time simultaneous interpretation at live events. Interpreters broadcast translate… | 77 | 1482 | active |
| k2-fsa/icefall Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi… | 64 | 1482 | active |
| IBM Watson Node.js SDK The official Node.js/TypeScript client library (npm package ibm-watson) for accessing IBM Watson services such as Assistant, Speech to Text… | 84 | 1473 | active |
| McCloudS/subgen Subgen is a self-hosted service that automatically generates subtitles for media files using OpenAI's Whisper speech recognition models. It… | 74 | 1465 | active |
| microsoft/NeuralSpeech NeuralSpeech is a Microsoft Research Asia research repository containing implementations of neural speech processing models across ASR erro… | 32 | 1462 | active |
| silverstein/minutes Minutes is an open-source, privacy-first conversation memory app that records meetings, voice memos, and dictation, transcribes them locall… | 77 | 1453 | active |
| DicioTeam/dicio-android Dicio is a free and open source voice assistant app for Android that interprets user questions and provides speech and graphical feedback e… | 83 | 1452 | active |
| sauravpanda/BrowserAI BrowserAI is a TypeScript library for running LLMs, speech recognition, text-to-speech, and audio separation models directly in the browser… | 78 | 1449 | active |
| jxlpzqc/TMSpeech TMSpeech is a Windows desktop application that captures system audio via WASAPI loopback and performs real-time Chinese speech-to-text, dis… | 82 | 1447 | active |
| edwko/OuteTTS OuteTTS is a Python interface for running OuteAI's text-to-speech models, supporting multiple inference backends including llama.cpp, Huggi… | 52 | 1436 | active |
| FireRedTeam/FireRedTTS2 FireRedTTS-2 is a long-form streaming text-to-speech system for multi-speaker dialogue generation, built in PyTorch with a dual-transformer… | 40 | 1428 | active |
| numz/sd-wav2lip-uhq A Wav2Lip Studio extension for the Stable Diffusion WebUI (Automatic1111) that generates high-quality lip-synced talking-face videos from a… | 28 | 1424 | active |
| rzru/nightingale Nightingale is an open-source karaoke application that turns any song in your music library into a karaoke track using neural networks. It … | 81 | 1423 | active |
| Voine/ChatWaifu_Mobile An Android app that lets users chat with an anime-style AI companion powered by ChatGPT, with local VITS text-to-speech, Live2D character r… | 64 | 1419 | active |
| lenML/Speech-AI-Forge Speech-AI-Forge is a Python project built around multiple TTS generation models (ChatTTS, CosyVoice, Fish-Speech, Index-TTS, F5-TTS, FireRe… | 62 | 1416 | active |
| R3gm/SoniTranslate SoniTranslate is a Gradio-based web application that automatically dubs videos into other languages. It transcribes speech, translates it, … | 55 | 1410 | active |
| joewongjc/type4me Type4Me is a macOS voice input method that transcribes speech in real time (streaming, local or cloud ASR engines) and refines text with LL… | 81 | 1404 | active |
| wxbool/video-srt-windows VideoSrt is an open-source Windows GUI tool written in Go that recognizes speech in video/audio files and automatically generates SRT subti… | 23 | 5032 | maintenance |
| 0nutation/SpeechGPT SpeechGPT is a series of speech large language models that can perceive and generate speech, supporting cross-modal instruction following a… | 29 | 1401 | active |
| elder-plinius/V3SP3R V3SP3R (Vesper) is an Android application that turns a Flipper Zero into an AI-controlled hardware hacking tool, driven by natural language… | 48 | 1400 | active |
| zackees/transcribe-anything A Python CLI app that transcribes local audio/video files or URLs (YouTube, Rumble, etc.) using multiple Whisper backends with automatic de… | 93 | 1393 | active |
| wenet-e2e/wespeaker WeSpeaker is a research and production-oriented toolkit for speaker embedding learning, supporting speaker verification, recognition, and d… | 64 | 1392 | active |
| 0xsline/OpenChatCut OpenChatCut is an open-source, local-first desktop AI video editor with a professional multi-track timeline that AI agents (built-in, Codex… | 80 | 1388 | active |
| callstackincubator/ai A collection of on-device AI primitives for React Native offering Vercel AI SDK-compatible APIs for running LLMs, embeddings, transcription… | 76 | 1387 | active |
| haoheliu/voicefixer VoiceFixer is a Python library and CLI tool for general speech restoration, using a pretrained neural vocoder to restore degraded human spe… | 26 | 1373 | stable |
| elevenyellow/handcrafted-persona-engine Persona Engine is a Windows desktop application that drives a Live2D avatar with an AI pipeline: microphone speech recognition, an LLM guid… | 73 | 1357 | active |
| k2-fsa/k2 k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)… | 64 | 1352 | active |
| nyrahealth/CrisperWhisper CrisperWhisper 2.0 is a controllable speech recognition model and Python library that transcribes audio either verbatim (including fillers,… | 89 | 1349 | active |
| Softcatala/whisper-ctranslate2 A command-line transcription and translation tool compatible with OpenAI's Whisper CLI, built on CTranslate2 and faster-whisper for up to 4… | 60 | 1343 | active |
| AI-FanGe/OpenAIglasses_for_Navigation An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st… | 39 | 1343 | active |
| claritylab/lucida Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros… | 32 | 4782 | maintenance |
| mbailey/voicemode VoiceMode is a Python-based MCP server and Claude Code plugin that enables natural, hands-free voice conversations with Claude Code and oth… | 80 | 1339 | active |
| mmorise/World WORLD is a C++ library for high-quality speech analysis, manipulation, and synthesis based on a vocoder design. It estimates F0 (via DIO/Ha… | 64 | 1338 | stable |
| google-gemma/gemma-translator An on-device, fully offline voice translator application powered by Gemma 4 via LiteRT-LM, with a React web frontend optimized for small ha… | 56 | 1337 | active |
| fayazara/Screendrop Screendrop is a native macOS app for screenshots and screen recording, serving as a free, self-hostable Loom alternative. It offers annotat… | 81 | 1333 | active |
| GauravSingh9356/J.A.R.V.I.S A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,… | 48 | 1332 | active |
| foobnix/LibreraReader Librera Reader is a highly customizable e-book reader application for Android supporting PDF, EPUB, MOBI, DjVu, FB2, CBZ/CBR and many other… | 97 | 4740 | maintenance |
| Nativ Nativ is a free, MIT-licensed macOS desktop application for running open AI models locally on Apple Silicon Macs, built on MLX-VLM. It prov… | 80 | 1330 | active |
| ratwithacompiler/OBS-captions-plugin An OBS Studio plugin that generates closed captions for live streams using the Google Cloud Speech Recognition API. It leverages Twitch's b… | 83 | 1324 | active |
| Robitx/gp.nvim Gp.nvim is a Neovim plugin that brings GPT-powered chat sessions, instructable text/code operations, speech-to-text, and image generation d… | 36 | 1320 | active |
| sanchit-gandhi/whisper-jax An optimized JAX implementation of OpenAI's Whisper speech recognition model, built on Hugging Face Transformers, offering up to 70x faster… | 30 | 4682 | maintenance |
| via007/bilibili-rag A self-hosted RAG application that turns a Bilibili user's favorited videos into a searchable, conversational knowledge base. It pulls favo… | 76 | 1306 | active |
| Henry-23/VideoChat A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late… | 48 | 1303 | active |
| stenolabs/stenoai Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire… | 84 | 1299 | active |
| espressif/esp-box ESP-BOX is Espressif's AIoT development framework for the ESP32-S3-BOX series of development boards, built on the ESP32-S3 Wi-Fi + Bluetoot… | 61 | 1297 | active |
| ictnlp/StreamSpeech StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p… | 37 | 1287 | active |
| AudileTeam/Audile Audile is an open-source Android app for music recognition that identifies songs playing nearby using AudD, ACRCloud, and Shazam services. … | 99 | 1284 | active |
| jasperproject/jasper-client Client code for the Jasper voice computing platform, an open source platform for building always-on, voice-controlled applications. It prov… | 32 | 4518 | maintenance |
| phuc-nt/my-translator A Tauri-based desktop app that captures system or microphone audio, transcribes it, and shows translations in a minimal real-time overlay, … | 76 | 1275 | active |
| ABexit/ASR-LLM-TTS An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) … | 60 | 1271 | active |
| jordanrendric/claude-video-vision A Claude Code plugin with an MCP server that gives Claude the ability to watch and understand videos by extracting frames via ffmpeg and tr… | 58 | 1265 | active |
| UltraStar-Deluxe/USDX UltraStar Deluxe is a free and open source karaoke singing game for PC, inspired by Sony's SingStar, where up to six (or twelve in party mo… | 99 | 1260 | active |
| yzfly/douyin-mcp-server A Python MCP server and WebUI that extracts watermark-free Douyin (TikTok China) video download links and transcribes video speech into tex… | 10 | 1258 | active |
| kubeai-project/kubeai KubeAI is a Kubernetes operator for serving machine learning models in production, supporting LLMs via vLLM and Ollama, vector embeddings, … | 89 | 1256 | active |
| peteonrails/voxtype Voxtype is a local, offline voice-to-text dictation tool for Linux with push-to-talk hotkey support across Wayland compositors like Hyprlan… | 78 | 1252 | active |
| FoloUp/FoloUp FoloUp is an open-source, self-hostable web application that conducts AI-powered voice interviews with job candidates. It generates tailore… | 58 | 1250 | active |
| Blueturboguy07/cue cue is an open-source desktop AI copilot that floats a glass panel over your screen, capturing your screen, microphone, and meeting audio t… | 78 | 1247 | active |
| sh-lee-prml/HierSpeechpp Official PyTorch implementation of HierSpeech++, a fast zero-shot speech synthesizer for text-to-speech and voice conversion based on hiera… | 28 | 1238 | active |
| OStudi/short-video-generator-AI An open-source Python tool that turns YouTube videos into ready-to-post vertical short videos by automatically detecting highlights, adding… | 57 | 1224 | active |
| yeyupiaoling/Whisper-Finetune A toolkit for fine-tuning OpenAI's Whisper speech recognition models using LoRA, supporting training with or without timestamps and even wi… | 66 | 1223 | active |
| SeaDve/Mousai Mousai is a GNOME desktop application written in Rust that identifies songs playing nearby, similar to Shazam. It listens via microphone or… | 80 | 1218 | active |
| tonyqinatcmu/SlideBot-AI SlideBot AI is an AI-powered presentation generator that turns a topic, outline, or uploaded materials (documents, spreadsheets, meeting re… | 44 | 1214 | active |
| uezo/ChatdollKit ChatdollKit is a Unity SDK that turns 3D character models into voice-enabled chatbots and virtual assistants. It integrates LLMs (ChatGPT, … | 76 | 1213 | active |
| modal-labs/quillman QuiLLMan is a voice chat application built on Kyutai's Moshi speech-to-speech language model, deployed serverlessly on Modal with a FastAPI… | 68 | 1213 | active |
| hgneng/ekho Ekho is an open-source Chinese text-to-speech engine supporting Mandarin, Cantonese, and Tibetan, part of the eGuideDog accessibility proje… | 68 | 1211 | active |
| ardha27/AI-Song-Cover-RVC A collection of Google Colab and Kaggle notebooks that form an all-in-one toolkit for creating AI song covers with RVC (Retrieval-based Voi… | 68 | 1207 | active |
| aTrainTranscription/aTrain aTrain is a desktop GUI application for offline transcription of speech recordings using Whisper-based machine learning models, with speake… | 79 | 1205 | active |
| hkjarral/AVA-AI-Voice-Agent-for-Asterisk An open-source AI voice agent that integrates with Asterisk/FreePBX phone systems via Audiosocket/RTP, built in Python with a modular pipel… | 85 | 1202 | active |
| nishuzumi/gemini-teacher A Python CLI application that acts as an English speaking practice assistant powered by Google Gemini. It listens to your speech via microp… | 68 | 1200 | active |
| kizuna-ai-lab/sokuji Sokuji is a cross-platform real-time two-way speech translation app for bilingual meetings, available as a desktop application (Windows, ma… | 83 | 1199 | active |
| aydinnyunus/ai-captcha-bypass A Python command-line tool that uses multimodal LLMs (GPT-4o, Gemini) to automatically solve various CAPTCHA types, including text, reCAPTC… | 61 | 1196 | active |
| diodiogod/TTS-Audio-Suite A ComfyUI custom node suite providing unified multi-engine Text-to-Speech, Voice Conversion, and audio editing across 19 engines like Chatt… | 84 | 1185 | active |
| digimata/parrot Parrot is an ultra-minimalist push-to-talk dictation daemon for macOS that transcribes speech entirely on-device using WhisperKit on the Ap… | 73 | 1185 | active |