function: speech-recognition
801 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| s-macke/SAM SAM (Software Automatic Mouth) is a tiny text-to-speech synthesizer written in C, adapted from the 1982 Commodore 64 speech software. It in… | 32 | 1497 | maintenance |
| wit-ai/pywit pywit is the official Python SDK for Wit.ai, Facebook's natural language processing platform. It provides a Wit client class for extracting… | 67 | 1485 | maintenance |
| fossasia/MMM-SUSI-AI A MagicMirror² module that integrates the SUSI.AI assistant, providing voice-activated intelligent answers on a smart mirror. It supports h… | 32 | 1482 | maintenance |
| fossasia/susi_alexa_skill An Amazon Alexa skill that connects Alexa-enabled devices to the Susi AI chatbot, letting users ask questions like 'Alexa, ask Susi what is… | 10 | 1470 | maintenance |
| microsoft/SpeechT5 Microsoft's unified-modal speech-text pre-training framework implementing SpeechT5 and related models (SpeechLM, SpeechUT, VALL-E X, WavLLM… | 32 | 1449 | maintenance |
| m1guelpf/yt-whisper A Python CLI tool that downloads YouTube videos with yt-dlp and generates subtitle files (VTT) using OpenAI's Whisper speech recognition mo… | 32 | 1446 | maintenance |
| cmusphinx/sphinx4 Sphinx-4 is a speaker-independent, continuous speech recognition library written entirely in Java, originating from CMU and industry resear… | 32 | 1437 | maintenance |
| hungtraan/FacebookBot A Facebook Messenger chatbot ('Optimist Prime') that supports voice recognition, natural language processing, and contextual follow-up conv… | 32 | 1424 | maintenance |
| DragonComputer/Dragonfire Dragonfire is an open-source virtual assistant for Ubuntu-based Linux distributions, combining speech recognition, text-to-speech, and NLP … | 23 | 1407 | maintenance |
| innnky/emotional-vits Emotional VITS is a voice synthesis (TTS) model based on VITS that enables emotion-controllable speech generation without requiring manual … | 32 | 1392 | maintenance |
| Jackywine/Bella Bella is a self-hosted Node.js web application that acts as a personalized AI digital companion with voice interaction. It combines Whisper… | 48 | 6379 | experimental |
| kripken/speak.js speak.js is a port of the eSpeak C++ speech synthesizer to JavaScript via Emscripten, enabling text-to-speech in the browser using only Jav… | 32 | 1335 | maintenance |
| CSTR-Edinburgh/merlin Merlin is a toolkit from the University of Edinburgh's CSTR for building deep neural network models for statistical parametric speech synth… | 23 | 1321 | maintenance |
| Renovamen/Speech-Emotion-Recognition A Python library implementing speech emotion recognition with Keras/TensorFlow 2 using LSTM, CNN, SVM, and MLP models. It extracts audio fe… | 32 | 1314 | maintenance |
| kakaobrain/pororo PORORO is a Python library from Kakao Brain that provides unified access to dozens of pretrained neural models for natural language process… | 10 | 1305 | maintenance |
| elanmart/cbp-translate A demo application that live-translates foreign-language speech in videos into subtitles, mimicking the Cyberpunk 2077 translation effect. … | 32 | 1274 | maintenance |
| sdkcarlos/artyom.js Artyom.js is a JavaScript library that wraps the Web Speech APIs (webkitSpeechRecognition and speechSynthesis) to add voice control, speech… | 23 | 1269 | maintenance |
| Alexander-H-Liu/End-to-end-ASR-Pytorch A PyTorch implementation of end-to-end automatic speech recognition (ASR), formerly known as Listen, Attend and Spell. It supports seq2seq … | 32 | 1208 | maintenance |
| Mentra-Community/OpenSourceSmartGlasses An open source smart glasses hardware and software project with display, microphones, wireless phone connection, and prescription lenses, i… | 23 | 1179 | maintenance |
| YaoFANGUK/video-subtitle-generator A Python application with both GUI and CLI interfaces that generates subtitle files (SRT) from video or audio using local Whisper-based spe… | 32 | 1175 | maintenance |
| openinterpreter/01 An open-source voice interface platform that lets users control computers conversationally, powered by Open Interpreter. It pairs a Python … | 16 | 5156 | experimental |
| RafalWilinski/telegram-chatgpt-concierge-bot A self-hosted Telegram bot that lets you chat with OpenAI's ChatGPT via text and voice messages. It uses LangChainJS for conversation histo… | 30 | 1130 | maintenance |
| synesthesiam/opentts OpenTTS is a text-to-speech server that unifies access to multiple open source TTS systems (Larynx, Glow-Speak, Coqui-TTS, MaryTTS, flite, … | 10 | 1118 | maintenance |
| synesthesiam/voice2json voice2json is a collection of command-line tools for offline speech-to-text and intent recognition on Linux, supporting 18 languages via en… | 10 | 1105 | maintenance |
| auspicious3000/autovc AUTOVC is a PyTorch implementation of a many-to-many non-parallel voice conversion framework that performs zero-shot voice style transfer u… | 32 | 1100 | maintenance |
| alumae/kaldi-gstreamer-server A real-time full-duplex speech recognition server built on the Kaldi toolkit and GStreamer framework, implemented in Python. It streams aud… | 32 | 1093 | maintenance |
| Yink/Amadeus An Android app that replicates the Amadeus AI assistant app from Steins;Gate 0, built primarily for cosplay purposes. It features speech re… | 23 | 1061 | maintenance |
| Edresson/YourTTS YourTTS is a zero-shot multi-speaker text-to-speech and voice conversion model built on VITS, implemented in the Coqui TTS framework. It su… | 23 | 1053 | maintenance |
| pykaldi/pykaldi PyKaldi is a Python scripting layer providing wrappers for the C++ APIs of the Kaldi speech recognition toolkit and OpenFst library. It ena… | 47 | 1039 | maintenance |
| ggeop/Python-ai-assistant Jarvis is a Python voice-controlled AI assistant for Linux that recognizes speech, responds conversationally, and executes commands like op… | 23 | 1014 | maintenance |
| cbh123/narrator A Python app that watches your webcam and generates David Attenborough-style narration of what it sees, using GPT vision models and ElevenL… | 69 | 4426 | experimental |
| susiai/susi_device SUSI Device provides sources to install the SUSI AI assistant stack on a Raspberry Pi, combining microphone/speaker, a small display, a loc… | 32 | 1004 | maintenance |
| Priler/jarvis JARVIS is an offline, privacy-respecting voice assistant built in Rust with Tauri, using neural networks for speech-to-text, text-to-speech… | 53 | 2909 | experimental |
| fikrikarim/parlor Parlor is a fully on-device, real-time multimodal voice assistant similar to GPT-Live, combining speech recognition, a Gemma vision-languag… | 78 | 2039 | experimental |
| everythingishacked/Semaphore Semaphore is a Python application that turns your full body into a keyboard using flag semaphore gestures. It uses OpenCV and MediaPipe pos… | 30 | 1939 | experimental |
| Mini-Omni Mini-Omni is an open-source multimodal large language model that performs real-time end-to-end speech-to-speech conversation with streaming… | 22 | 1922 | experimental |
| Standard-Intelligence/hertz-dev Hertz-dev is an open-source 8.5B parameter autoregressive base model for full-duplex conversational audio, released by Standard Intelligenc… | 22 | 1799 | experimental |
| antirez/voxtral.c A pure C, zero-dependency inference implementation of Mistral's Voxtral Realtime 4B speech-to-text model, with a streaming C API and CLI fo… | 45 | 1738 | experimental |
| collabora/WhisperFusion WhisperFusion is a real-time voice chat application that combines WhisperLive speech-to-text, a Mistral/Phi LLM, and WhisperSpeech text-to-… | 26 | 1647 | experimental |
| linyiLYi/voice-assistant A simple single-script Python demo of a local voice assistant that uses Whisper (via Apple MLX) for speech recognition and a local Yi large… | 26 | 1323 | experimental |
| elfvingralf/macOSpilot-ai-assistant macOSpilot is a macOS desktop AI assistant built with Electron that answers spoken or typed questions about whatever application is current… | 26 | 1157 | experimental |
| CSCB/vibe-mouse An open-source Python desktop application that redefines mouse interaction by letting users bind customizable 'Skills' to buttons, with mul… | 55 | 1086 | experimental |
| read-cat/read-cat ReadCat is a free, open-source, ad-free novel reader built with TypeScript. It supports online book sources via plugins, local txt books, b… | 22 | 1010 | experimental |
| ronibandini/reggaetonBeGone A Raspberry Pi-based edge machine learning device that continuously samples ambient audio and uses an Edge Impulse audio classification mod… | 70 | 1005 | experimental |
| susiai/susi_shell A suite of Python-based command line tools for interacting with AI services directly from the terminal, including chat, text completion, tr… | 41 | 1002 | experimental |
| mozilla/DeepSpeech DeepSpeech is an open-source, offline speech-to-text engine based on Baidu's Deep Speech research paper and implemented with TensorFlow. It… | 10 | 26771 | abandoned |
| supertone-inc/supertonic Supertonic is a lightning-fast, on-device multilingual text-to-speech system powered by ONNX Runtime, with a compact 99M-parameter open-wei… | 58 | 13734 | abandoned |
| fuergaosi233/wechat-chatgpt A TypeScript bot that connects OpenAI's ChatGPT API to WeChat using the wechaty library, supporting text conversations, DALL·E image genera… | 22 | 13238 | abandoned |
| voicepaw/so-vits-svc-fork A fork of so-vits-svc providing singing voice conversion with realtime support and an improved interface, built on PyTorch and PyTorch Ligh… | 82 | 9327 | abandoned |
| MycroftAI/mycroft-core Mycroft Core is the core software of the Mycroft open-source voice assistant platform, providing wake-word listening, speech recognition, n… | 10 | 6611 | abandoned |
| unCaptcha unCaptcha2 is a Python-based security research tool that defeats Google's ReCaptcha v2 audio challenges by submitting the audio to free spe… | 32 | 4917 | abandoned |
| agermanidis/autosub Autosub is a Python command-line utility that auto-generates subtitles for video or audio files. It performs voice activity detection, tran… | 32 | 4191 | abandoned |
| buriburisuri/speech-to-text-wavenet A TensorFlow implementation of end-to-end English speech recognition based on DeepMind's WaveNet architecture, trained with CTC loss on sen… | 32 | 4002 | abandoned |
| innnky/so-vits-svc A singing voice conversion (SVC) framework that uses a SoftVC content encoder with VITS to transform one singer's voice into another timbre… | 10 | 3779 | abandoned |
| askrella/whatsapp-chatgpt A self-hosted WhatsApp bot that connects OpenAI's GPT and DALL-E 2 to WhatsApp, letting users chat with an AI assistant and generate images… | 73 | 3777 | abandoned |
| adamcohenhillel/ADeus Adeus is an open-source AI wearable project that records what you say and hear, transcribes it, and stores it on your own server (Supabase … | 26 | 3426 | abandoned |
| Kitt-AI/snowboy Snowboy is a C++-based hotword (wake word) detection library by KITT.AI with bindings for Python, Android, and other platforms, enabling al… | 23 | 3364 | abandoned |
| zzw922cn/Automatic_Speech_Recognition An end-to-end automatic speech recognition system implemented in TensorFlow, supporting Mandarin and English with models like DeepSpeech2, … | 32 | 2831 | abandoned |
| coqui-ai/STT Coqui STT is an open-source deep learning toolkit for training and deploying speech-to-text models, built on TensorFlow with bindings for m… | 23 | 2606 | abandoned |
| idootop/open-xiaoai Open-XiaoAI is a Rust-based client/server project that takes over the audio input and output of Xiaomi XiaoAI smart speakers (LX06 and OH2P… | 10 | 2595 | abandoned |
| react-native-voice/voice A React Native speech-to-text library providing voice recognition on iOS and Android with both online and offline support. The package is n… | 10 | 2162 | abandoned |
| joshnewlan/say_what A Python script that listens to conference call audio via speech-to-text (IBM Watson) and alerts the user on HipChat when their name is men… | 32 | 2087 | abandoned |
| C-Nedelcu/talk-to-chatgpt A Chrome and Edge browser extension that lets users talk to ChatGPT using speech recognition and hear responses via text-to-speech, with op… | 32 | 1929 | abandoned |
| wzpan/dingdang-robot Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi. It uses pluggable STT/TTS en… | 10 | 1873 | abandoned |
| justLV/onju-voice A hackable AI home assistant platform that replaces the internals of a Google Nest Mini with a custom ESP32-S3 PCB, paired with a server th… | 63 | 1711 | abandoned |
| Ayanaminn/N46Whisper A Google Colab notebook application that generates Japanese subtitle files from video using the faster-whisper speech recognition model. It… | 36 | 1708 | abandoned |
| NVIDIA/OpenSeq2Seq OpenSeq2Seq is a TensorFlow-based toolkit for building and training sequence-to-sequence models for neural machine translation, speech reco… | 10 | 1558 | abandoned |
| elevenlabs/elevenlabs-mcp The official ElevenLabs Model Context Protocol (MCP) server, exposing ElevenLabs text-to-speech, speech-to-text, voice design, and conversa… | 10 | 1532 | abandoned |
| fossasia/susi_smart_box SUSI.AI Smart Box is an open-source smart speaker / voice assistant hardware project built around the SUSI.AI assistant. It provides the so… | 10 | 1529 | abandoned |
| fossasia/susi_desktop An Electron-based desktop client for the SUSI AI open-source personal assistant, connecting to the api.susi.ai server. It supports chat and… | 10 | 1508 | abandoned |
| adblockradio/adblockradio Adblock Radio is a Node.js library that detects and blocks advertisements in live radio streams and podcasts using machine learning and aud… | 10 | 1488 | abandoned |
| sc0ty/subsync A C++ tool that automatically synchronizes subtitle files with movie or TV audio using speech recognition. It detects spoken audio in the v… | 10 | 1423 | abandoned |
| zenorocha/voice-elements A pair of Polymer-based Web Components (<voice-player> and <voice-recognition>) that wrap the Web Speech API for speech synthesis (text to … | 23 | 1348 | abandoned |
| dingdang-robot/dingdang-robot Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi and other Linux hosts. It use… | 10 | 1330 | abandoned |
| alexa-pi/AlexaPi AlexaPi is an open-source Python client for Amazon's Alexa voice service designed to run on devices like Raspberry Pi, Orange Pi, CHIP, and… | 10 | 1327 | abandoned |
| jiran214/GPT-vup GPT-vup is a Python application that powers an AI-driven Live2D virtual streamer (VTuber) on BiliBili and Douyin live platforms. It uses Op… | 30 | 1271 | abandoned |
| MycroftAI/mimic3 Mimic 3 is a fast, local neural text-to-speech engine developed by Mycroft for the Mark II voice assistant, usable as a Python library, CLI… | 29 | 1264 | abandoned |
| alexa/avs-device-sdk The Alexa Voice Service (AVS) Device SDK is a C++ SDK for commercial device makers to integrate Alexa voice assistant capabilities directly… | 10 | 1252 | abandoned |
| rhasspy/wyoming-satellite A Python application that turns a Raspberry Pi (or similar Linux device) with a microphone and speaker into a remote voice satellite using … | 10 | 1244 | abandoned |
| yodaos-project/yodaos YodaOS is a Linux distribution built on OpenWrt for voice-enabled IoT devices, using JavaScript as its primary application language. It tar… | 32 | 1225 | abandoned |
| MicrosoftEdge/magic-mirror-demo A smart mirror IoT demo project by Microsoft Edge that displays information on a two-way mirror and recognizes registered users via facial … | 10 | 1225 | abandoned |
| xiph/LPCNet LPCNet is a low-complexity C implementation of the WaveRNN-based LPCNet neural vocoder for efficient speech synthesis and compression. It a… | 32 | 1221 | abandoned |
| amzn/alexa-skills-kit-js The original Node.js SDK and example code for building voice-enabled Alexa skills for Amazon Echo devices. This repository is deprecated an… | 10 | 1146 | abandoned |
| shivasiddharth/GassistPi GassistPi is a Python application that turns single board computers like the Raspberry Pi into a Google Assistant-powered voice assistant. … | 44 | 1031 | abandoned |
| BrasD99/HeyGenClone An open-source Python application that clones the HeyGen video translation system, translating videos into multiple languages with voice ov… | 10 | 1027 | abandoned |
| huggingface/transformers Hugging Face Transformers is a Python library that serves as the model-definition framework for state-of-the-art machine learning models ac… | 95 | 164475 | stable |
| Mintplex-Labs/anything-llm AnythingLLM is an all-in-one, local-first AI application for chatting with your documents, running AI agents, and building workflows entire… | 95 | 65257 | active |
| mudler/LocalAI LocalAI is an open-source, self-hosted AI inference runtime that runs LLMs, vision, speech, image, and video models behind OpenAI/Anthropic… | 93 | 48696 | active |
| khoj-ai/khoj Khoj is a self-hostable personal AI assistant ('second brain') that answers questions from your own documents and the web using local or on… | 85 | 36730 | active |
| 78/xiaozhi-esp32 XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen… | 88 | 29186 | active |
| openai/openai-agents-python OpenAI Agents SDK is a lightweight Python framework for building multi-agent LLM workflows with agents, handoffs, tools, guardrails, sessio… | 82 | 28982 | active |
| MLX MLX is an array computation framework for machine learning on Apple silicon, developed by Apple ML research. It offers NumPy-like Python AP… | 94 | 28172 | active |
| Fosowl/agenticSeek AgenticSeek is a fully local, privacy-focused AI assistant that serves as an open-source alternative to Manus AI. It runs reasoning models … | 64 | 27027 | active |
| OpenBMB/MiniCPM-V MiniCPM-V and MiniCPM-o are a series of small multimodal large language models for efficient image, video, and audio understanding, deploya… | 61 | 26240 | active |
| screenpipe/screenpipe Screenpipe is a source-available desktop application that continuously records your screen and audio locally, extracting text via OCR/acces… | 86 | 21244 | active |
| huggingface/candle Candle is a minimalist machine learning framework for Rust focused on performance and ease of use, with CPU and CUDA GPU support. It ships … | 73 | 20955 | active |
| THU-MAIC/OpenMAIC OpenMAIC is an open-source multi-agent interactive classroom that delivers immersive AI-driven learning experiences with one click. Built w… | 81 | 20934 | active |
| livekit/livekit LiveKit is an open-source, scalable WebRTC SFU media server written in Go that provides realtime video, audio, and data transport for appli… | 99 | 20530 | stable |
| pydantic/pydantic-ai Pydantic AI is a type-safe Python SDK for building AI agents and LLM applications, with a typed agent loop, structured outputs, tool callin… | 86 | 19518 | active |
| arc53/DocsGPT DocsGPT is an open-source AI platform for building private agents, assistants, and enterprise search over your own documents. It includes a… | 93 | 18229 | active |