Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: speech-recognition

801 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
s-macke/SAM
SAM (Software Automatic Mouth) is a tiny text-to-speech synthesizer written in C, adapted from the 1982 Commodore 64 speech software. It in…
321497maintenance
wit-ai/pywit
pywit is the official Python SDK for Wit.ai, Facebook's natural language processing platform. It provides a Wit client class for extracting…
671485maintenance
fossasia/MMM-SUSI-AI
A MagicMirror² module that integrates the SUSI.AI assistant, providing voice-activated intelligent answers on a smart mirror. It supports h…
321482maintenance
fossasia/susi_alexa_skill
An Amazon Alexa skill that connects Alexa-enabled devices to the Susi AI chatbot, letting users ask questions like 'Alexa, ask Susi what is…
101470maintenance
microsoft/SpeechT5
Microsoft's unified-modal speech-text pre-training framework implementing SpeechT5 and related models (SpeechLM, SpeechUT, VALL-E X, WavLLM…
321449maintenance
m1guelpf/yt-whisper
A Python CLI tool that downloads YouTube videos with yt-dlp and generates subtitle files (VTT) using OpenAI's Whisper speech recognition mo…
321446maintenance
cmusphinx/sphinx4
Sphinx-4 is a speaker-independent, continuous speech recognition library written entirely in Java, originating from CMU and industry resear…
321437maintenance
hungtraan/FacebookBot
A Facebook Messenger chatbot ('Optimist Prime') that supports voice recognition, natural language processing, and contextual follow-up conv…
321424maintenance
DragonComputer/Dragonfire
Dragonfire is an open-source virtual assistant for Ubuntu-based Linux distributions, combining speech recognition, text-to-speech, and NLP …
231407maintenance
innnky/emotional-vits
Emotional VITS is a voice synthesis (TTS) model based on VITS that enables emotion-controllable speech generation without requiring manual …
321392maintenance
Jackywine/Bella
Bella is a self-hosted Node.js web application that acts as a personalized AI digital companion with voice interaction. It combines Whisper…
486379experimental
kripken/speak.js
speak.js is a port of the eSpeak C++ speech synthesizer to JavaScript via Emscripten, enabling text-to-speech in the browser using only Jav…
321335maintenance
CSTR-Edinburgh/merlin
Merlin is a toolkit from the University of Edinburgh's CSTR for building deep neural network models for statistical parametric speech synth…
231321maintenance
Renovamen/Speech-Emotion-Recognition
A Python library implementing speech emotion recognition with Keras/TensorFlow 2 using LSTM, CNN, SVM, and MLP models. It extracts audio fe…
321314maintenance
kakaobrain/pororo
PORORO is a Python library from Kakao Brain that provides unified access to dozens of pretrained neural models for natural language process…
101305maintenance
elanmart/cbp-translate
A demo application that live-translates foreign-language speech in videos into subtitles, mimicking the Cyberpunk 2077 translation effect. …
321274maintenance
sdkcarlos/artyom.js
Artyom.js is a JavaScript library that wraps the Web Speech APIs (webkitSpeechRecognition and speechSynthesis) to add voice control, speech…
231269maintenance
Alexander-H-Liu/End-to-end-ASR-Pytorch
A PyTorch implementation of end-to-end automatic speech recognition (ASR), formerly known as Listen, Attend and Spell. It supports seq2seq …
321208maintenance
Mentra-Community/OpenSourceSmartGlasses
An open source smart glasses hardware and software project with display, microphones, wireless phone connection, and prescription lenses, i…
231179maintenance
YaoFANGUK/video-subtitle-generator
A Python application with both GUI and CLI interfaces that generates subtitle files (SRT) from video or audio using local Whisper-based spe…
321175maintenance
openinterpreter/01
An open-source voice interface platform that lets users control computers conversationally, powered by Open Interpreter. It pairs a Python …
165156experimental
RafalWilinski/telegram-chatgpt-concierge-bot
A self-hosted Telegram bot that lets you chat with OpenAI's ChatGPT via text and voice messages. It uses LangChainJS for conversation histo…
301130maintenance
synesthesiam/opentts
OpenTTS is a text-to-speech server that unifies access to multiple open source TTS systems (Larynx, Glow-Speak, Coqui-TTS, MaryTTS, flite, …
101118maintenance
synesthesiam/voice2json
voice2json is a collection of command-line tools for offline speech-to-text and intent recognition on Linux, supporting 18 languages via en…
101105maintenance
auspicious3000/autovc
AUTOVC is a PyTorch implementation of a many-to-many non-parallel voice conversion framework that performs zero-shot voice style transfer u…
321100maintenance
alumae/kaldi-gstreamer-server
A real-time full-duplex speech recognition server built on the Kaldi toolkit and GStreamer framework, implemented in Python. It streams aud…
321093maintenance
Yink/Amadeus
An Android app that replicates the Amadeus AI assistant app from Steins;Gate 0, built primarily for cosplay purposes. It features speech re…
231061maintenance
Edresson/YourTTS
YourTTS is a zero-shot multi-speaker text-to-speech and voice conversion model built on VITS, implemented in the Coqui TTS framework. It su…
231053maintenance
pykaldi/pykaldi
PyKaldi is a Python scripting layer providing wrappers for the C++ APIs of the Kaldi speech recognition toolkit and OpenFst library. It ena…
471039maintenance
ggeop/Python-ai-assistant
Jarvis is a Python voice-controlled AI assistant for Linux that recognizes speech, responds conversationally, and executes commands like op…
231014maintenance
cbh123/narrator
A Python app that watches your webcam and generates David Attenborough-style narration of what it sees, using GPT vision models and ElevenL…
694426experimental
susiai/susi_device
SUSI Device provides sources to install the SUSI AI assistant stack on a Raspberry Pi, combining microphone/speaker, a small display, a loc…
321004maintenance
Priler/jarvis
JARVIS is an offline, privacy-respecting voice assistant built in Rust with Tauri, using neural networks for speech-to-text, text-to-speech…
532909experimental
fikrikarim/parlor
Parlor is a fully on-device, real-time multimodal voice assistant similar to GPT-Live, combining speech recognition, a Gemma vision-languag…
782039experimental
everythingishacked/Semaphore
Semaphore is a Python application that turns your full body into a keyboard using flag semaphore gestures. It uses OpenCV and MediaPipe pos…
301939experimental
Mini-Omni
Mini-Omni is an open-source multimodal large language model that performs real-time end-to-end speech-to-speech conversation with streaming…
221922experimental
Standard-Intelligence/hertz-dev
Hertz-dev is an open-source 8.5B parameter autoregressive base model for full-duplex conversational audio, released by Standard Intelligenc…
221799experimental
antirez/voxtral.c
A pure C, zero-dependency inference implementation of Mistral's Voxtral Realtime 4B speech-to-text model, with a streaming C API and CLI fo…
451738experimental
collabora/WhisperFusion
WhisperFusion is a real-time voice chat application that combines WhisperLive speech-to-text, a Mistral/Phi LLM, and WhisperSpeech text-to-…
261647experimental
linyiLYi/voice-assistant
A simple single-script Python demo of a local voice assistant that uses Whisper (via Apple MLX) for speech recognition and a local Yi large…
261323experimental
elfvingralf/macOSpilot-ai-assistant
macOSpilot is a macOS desktop AI assistant built with Electron that answers spoken or typed questions about whatever application is current…
261157experimental
CSCB/vibe-mouse
An open-source Python desktop application that redefines mouse interaction by letting users bind customizable 'Skills' to buttons, with mul…
551086experimental
read-cat/read-cat
ReadCat is a free, open-source, ad-free novel reader built with TypeScript. It supports online book sources via plugins, local txt books, b…
221010experimental
ronibandini/reggaetonBeGone
A Raspberry Pi-based edge machine learning device that continuously samples ambient audio and uses an Edge Impulse audio classification mod…
701005experimental
susiai/susi_shell
A suite of Python-based command line tools for interacting with AI services directly from the terminal, including chat, text completion, tr…
411002experimental
mozilla/DeepSpeech
DeepSpeech is an open-source, offline speech-to-text engine based on Baidu's Deep Speech research paper and implemented with TensorFlow. It…
1026771abandoned
supertone-inc/supertonic
Supertonic is a lightning-fast, on-device multilingual text-to-speech system powered by ONNX Runtime, with a compact 99M-parameter open-wei…
5813734abandoned
fuergaosi233/wechat-chatgpt
A TypeScript bot that connects OpenAI's ChatGPT API to WeChat using the wechaty library, supporting text conversations, DALL·E image genera…
2213238abandoned
voicepaw/so-vits-svc-fork
A fork of so-vits-svc providing singing voice conversion with realtime support and an improved interface, built on PyTorch and PyTorch Ligh…
829327abandoned
MycroftAI/mycroft-core
Mycroft Core is the core software of the Mycroft open-source voice assistant platform, providing wake-word listening, speech recognition, n…
106611abandoned
unCaptcha
unCaptcha2 is a Python-based security research tool that defeats Google's ReCaptcha v2 audio challenges by submitting the audio to free spe…
324917abandoned
agermanidis/autosub
Autosub is a Python command-line utility that auto-generates subtitles for video or audio files. It performs voice activity detection, tran…
324191abandoned
buriburisuri/speech-to-text-wavenet
A TensorFlow implementation of end-to-end English speech recognition based on DeepMind's WaveNet architecture, trained with CTC loss on sen…
324002abandoned
innnky/so-vits-svc
A singing voice conversion (SVC) framework that uses a SoftVC content encoder with VITS to transform one singer's voice into another timbre…
103779abandoned
askrella/whatsapp-chatgpt
A self-hosted WhatsApp bot that connects OpenAI's GPT and DALL-E 2 to WhatsApp, letting users chat with an AI assistant and generate images…
733777abandoned
adamcohenhillel/ADeus
Adeus is an open-source AI wearable project that records what you say and hear, transcribes it, and stores it on your own server (Supabase …
263426abandoned
Kitt-AI/snowboy
Snowboy is a C++-based hotword (wake word) detection library by KITT.AI with bindings for Python, Android, and other platforms, enabling al…
233364abandoned
zzw922cn/Automatic_Speech_Recognition
An end-to-end automatic speech recognition system implemented in TensorFlow, supporting Mandarin and English with models like DeepSpeech2, …
322831abandoned
coqui-ai/STT
Coqui STT is an open-source deep learning toolkit for training and deploying speech-to-text models, built on TensorFlow with bindings for m…
232606abandoned
idootop/open-xiaoai
Open-XiaoAI is a Rust-based client/server project that takes over the audio input and output of Xiaomi XiaoAI smart speakers (LX06 and OH2P…
102595abandoned
react-native-voice/voice
A React Native speech-to-text library providing voice recognition on iOS and Android with both online and offline support. The package is n…
102162abandoned
joshnewlan/say_what
A Python script that listens to conference call audio via speech-to-text (IBM Watson) and alerts the user on HipChat when their name is men…
322087abandoned
C-Nedelcu/talk-to-chatgpt
A Chrome and Edge browser extension that lets users talk to ChatGPT using speech recognition and hear responses via text-to-speech, with op…
321929abandoned
wzpan/dingdang-robot
Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi. It uses pluggable STT/TTS en…
101873abandoned
justLV/onju-voice
A hackable AI home assistant platform that replaces the internals of a Google Nest Mini with a custom ESP32-S3 PCB, paired with a server th…
631711abandoned
Ayanaminn/N46Whisper
A Google Colab notebook application that generates Japanese subtitle files from video using the faster-whisper speech recognition model. It…
361708abandoned
NVIDIA/OpenSeq2Seq
OpenSeq2Seq is a TensorFlow-based toolkit for building and training sequence-to-sequence models for neural machine translation, speech reco…
101558abandoned
elevenlabs/elevenlabs-mcp
The official ElevenLabs Model Context Protocol (MCP) server, exposing ElevenLabs text-to-speech, speech-to-text, voice design, and conversa…
101532abandoned
fossasia/susi_smart_box
SUSI.AI Smart Box is an open-source smart speaker / voice assistant hardware project built around the SUSI.AI assistant. It provides the so…
101529abandoned
fossasia/susi_desktop
An Electron-based desktop client for the SUSI AI open-source personal assistant, connecting to the api.susi.ai server. It supports chat and…
101508abandoned
adblockradio/adblockradio
Adblock Radio is a Node.js library that detects and blocks advertisements in live radio streams and podcasts using machine learning and aud…
101488abandoned
sc0ty/subsync
A C++ tool that automatically synchronizes subtitle files with movie or TV audio using speech recognition. It detects spoken audio in the v…
101423abandoned
zenorocha/voice-elements
A pair of Polymer-based Web Components (<voice-player> and <voice-recognition>) that wrap the Web Speech API for speech synthesis (text to …
231348abandoned
dingdang-robot/dingdang-robot
Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi and other Linux hosts. It use…
101330abandoned
alexa-pi/AlexaPi
AlexaPi is an open-source Python client for Amazon's Alexa voice service designed to run on devices like Raspberry Pi, Orange Pi, CHIP, and…
101327abandoned
jiran214/GPT-vup
GPT-vup is a Python application that powers an AI-driven Live2D virtual streamer (VTuber) on BiliBili and Douyin live platforms. It uses Op…
301271abandoned
MycroftAI/mimic3
Mimic 3 is a fast, local neural text-to-speech engine developed by Mycroft for the Mark II voice assistant, usable as a Python library, CLI…
291264abandoned
alexa/avs-device-sdk
The Alexa Voice Service (AVS) Device SDK is a C++ SDK for commercial device makers to integrate Alexa voice assistant capabilities directly…
101252abandoned
rhasspy/wyoming-satellite
A Python application that turns a Raspberry Pi (or similar Linux device) with a microphone and speaker into a remote voice satellite using …
101244abandoned
yodaos-project/yodaos
YodaOS is a Linux distribution built on OpenWrt for voice-enabled IoT devices, using JavaScript as its primary application language. It tar…
321225abandoned
MicrosoftEdge/magic-mirror-demo
A smart mirror IoT demo project by Microsoft Edge that displays information on a two-way mirror and recognizes registered users via facial …
101225abandoned
xiph/LPCNet
LPCNet is a low-complexity C implementation of the WaveRNN-based LPCNet neural vocoder for efficient speech synthesis and compression. It a…
321221abandoned
amzn/alexa-skills-kit-js
The original Node.js SDK and example code for building voice-enabled Alexa skills for Amazon Echo devices. This repository is deprecated an…
101146abandoned
shivasiddharth/GassistPi
GassistPi is a Python application that turns single board computers like the Raspberry Pi into a Google Assistant-powered voice assistant. …
441031abandoned
BrasD99/HeyGenClone
An open-source Python application that clones the HeyGen video translation system, translating videos into multiple languages with voice ov…
101027abandoned
huggingface/transformers
Hugging Face Transformers is a Python library that serves as the model-definition framework for state-of-the-art machine learning models ac…
95164475stable
Mintplex-Labs/anything-llm
AnythingLLM is an all-in-one, local-first AI application for chatting with your documents, running AI agents, and building workflows entire…
9565257active
mudler/LocalAI
LocalAI is an open-source, self-hosted AI inference runtime that runs LLMs, vision, speech, image, and video models behind OpenAI/Anthropic…
9348696active
khoj-ai/khoj
Khoj is a self-hostable personal AI assistant ('second brain') that answers questions from your own documents and the web using local or on…
8536730active
78/xiaozhi-esp32
XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen…
8829186active
openai/openai-agents-python
OpenAI Agents SDK is a lightweight Python framework for building multi-agent LLM workflows with agents, handoffs, tools, guardrails, sessio…
8228982active
MLX
MLX is an array computation framework for machine learning on Apple silicon, developed by Apple ML research. It offers NumPy-like Python AP…
9428172active
Fosowl/agenticSeek
AgenticSeek is a fully local, privacy-focused AI assistant that serves as an open-source alternative to Manus AI. It runs reasoning models …
6427027active
OpenBMB/MiniCPM-V
MiniCPM-V and MiniCPM-o are a series of small multimodal large language models for efficient image, video, and audio understanding, deploya…
6126240active
screenpipe/screenpipe
Screenpipe is a source-available desktop application that continuously records your screen and audio locally, extracting text via OCR/acces…
8621244active
huggingface/candle
Candle is a minimalist machine learning framework for Rust focused on performance and ease of use, with CPU and CUDA GPU support. It ships …
7320955active
THU-MAIC/OpenMAIC
OpenMAIC is an open-source multi-agent interactive classroom that delivers immersive AI-driven learning experiences with one click. Built w…
8120934active
livekit/livekit
LiveKit is an open-source, scalable WebRTC SFU media server written in Go that provides realtime video, audio, and data transport for appli…
9920530stable
pydantic/pydantic-ai
Pydantic AI is a type-safe Python SDK for building AI agents and LLM applications, with a typed agent loop, structured outputs, tool callin…
8619518active
arc53/DocsGPT
DocsGPT is an open-source AI platform for building private agents, assistants, and enterprise search over your own documents. It includes a…
9318229active

← prev page 6 / 9 next →