Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: speech-recognition

801 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Lynpoint/CyberVerse
CyberVerse is an open-source, self-hosted framework for building real-time, voice-first AI agents with optional digital-human video (talkin…
631612active
mkiol/dsnote
Speech Note is a Linux desktop and Sailfish OS application for note taking, reading, and translating text using offline Speech to Text, Tex…
961607active
modem-works/dream-recorder
Dream Recorder is an open-source DIY bedside device built on a Raspberry Pi 5 that records users speaking their dreams aloud and turns them…
411607active
Roy3838/Observer
Observer AI is a desktop application for building micro-agents that observe screen, camera, microphone, and audio inputs, process them with…
871601active
Norsico/Video-Materials-AutoGEN-Workstation
A self-hosted short-video production workstation that combines AI script generation (Gemini), batch TTS voiceover, AI image asset synthesis…
541599active
SAGAR-TAMANG/friday-tony-stark-demo
F.R.I.D.A.Y. is a Tony Stark-inspired voice AI assistant consisting of a FastMCP server exposing tools (web search, news, system info) and …
551595active
semperai/amica
Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. …
321593active
finnvoor/yap
yap is a Swift CLI for on-device speech transcription of audio and video files using Apple's Speech.framework on macOS 26. It also supports…
811586active
royshil/obs-localvocal
LocalVocal is an OBS Studio plugin that performs real-time, fully local speech recognition and translation using Whisper models running on …
811585active
TOM88812/xiaozhi-android-client
A Flutter-based cross-platform voice chat client for the xiaozhi AI assistant ecosystem, supporting real-time voice interaction and text co…
601574active
bootphon/phonemizer
A Python library and command-line tool that converts words and texts into phonemes across many languages. It wraps multiple backends (espea…
851567stable
byjlw/video-analyzer
A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs …
601559active
RunanywhereAI/RCLI
RCLI is a single-binary CLI that runs open-source AI models locally on your machine, covering chat, vision, speech-to-text, text-to-speech,…
821542active
vibevoice-community/VibeVoice
VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker …
601542active
dindin0497/SeeIt
SeeIt is an inclusive Android app with two accessibility modes: one that converts spoken speech into text and plays corresponding ASL (Amer…
411540active
tencent-connect/openclaw-qqbot
A message channel plugin for OpenClaw that connects an AI assistant to QQ (Tencent's messaging platform), supporting private chat, group ch…
821538active
pipecat-ai/smart-turn
Smart Turn is an open-source, audio-native turn detection model that decides when a voice agent should respond to human speech, using proso…
491538active
AlekPet/ComfyUI_Custom_Nodes_AlekPet
A collection of custom nodes for ComfyUI that extend its capabilities with painting, pose control, prompt translation, and speech recogniti…
721524active
bytedance/SALMONN
SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video…
731513active
kyutai-labs/hibiki
Hibiki is a decoder-only model for streaming (simultaneous) speech-to-speech translation, built on the multistream Moshi architecture. It p…
281509active
amicalhq/amical
Amical is an open-source, local-first AI dictation app for macOS and Windows that transcribes speech with Whisper and enhances it with LLMs…
841506active
kyutai-labs/unmute
Unmute is a system that lets any text LLM listen and speak by wrapping it with Kyutai's low-latency speech-to-text and text-to-speech model…
601506active
Flashlight
wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into …
625466maintenance
stepfun-ai/Step-Audio2
Step-Audio 2 is an end-to-end multimodal large language model for industry-strength audio understanding and speech conversation, with open-…
501503active
QwenAudio/Fun-ASR
Fun-ASR is a family of open-source LLM-based end-to-end speech recognition models from Tongyi Lab, covering Chinese, dialects, accents, and…
811496active
espressif/esp-sr
ESP-SR is Espressif's speech recognition framework for ESP32-series chips, providing wake word detection (WakeNet), voice activity detectio…
781492active
ekwek1/soprano
Soprano is an ultra-lightweight 80M-parameter text-to-speech model and Python library for fast, expressive, high-fidelity speech synthesis …
441486active
fossasia/voxbento
Voxbento is an open-source, self-hosted platform for real-time simultaneous interpretation at live events. Interpreters broadcast translate…
771482active
k2-fsa/icefall
Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi…
641482active
IBM Watson Node.js SDK
The official Node.js/TypeScript client library (npm package ibm-watson) for accessing IBM Watson services such as Assistant, Speech to Text…
841473active
McCloudS/subgen
Subgen is a self-hosted service that automatically generates subtitles for media files using OpenAI's Whisper speech recognition models. It…
741465active
microsoft/NeuralSpeech
NeuralSpeech is a Microsoft Research Asia research repository containing implementations of neural speech processing models across ASR erro…
321462active
silverstein/minutes
Minutes is an open-source, privacy-first conversation memory app that records meetings, voice memos, and dictation, transcribes them locall…
771453active
DicioTeam/dicio-android
Dicio is a free and open source voice assistant app for Android that interprets user questions and provides speech and graphical feedback e…
831452active
sauravpanda/BrowserAI
BrowserAI is a TypeScript library for running LLMs, speech recognition, text-to-speech, and audio separation models directly in the browser…
781449active
jxlpzqc/TMSpeech
TMSpeech is a Windows desktop application that captures system audio via WASAPI loopback and performs real-time Chinese speech-to-text, dis…
821447active
edwko/OuteTTS
OuteTTS is a Python interface for running OuteAI's text-to-speech models, supporting multiple inference backends including llama.cpp, Huggi…
521436active
FireRedTeam/FireRedTTS2
FireRedTTS-2 is a long-form streaming text-to-speech system for multi-speaker dialogue generation, built in PyTorch with a dual-transformer…
401428active
numz/sd-wav2lip-uhq
A Wav2Lip Studio extension for the Stable Diffusion WebUI (Automatic1111) that generates high-quality lip-synced talking-face videos from a…
281424active
rzru/nightingale
Nightingale is an open-source karaoke application that turns any song in your music library into a karaoke track using neural networks. It …
811423active
Voine/ChatWaifu_Mobile
An Android app that lets users chat with an anime-style AI companion powered by ChatGPT, with local VITS text-to-speech, Live2D character r…
641419active
lenML/Speech-AI-Forge
Speech-AI-Forge is a Python project built around multiple TTS generation models (ChatTTS, CosyVoice, Fish-Speech, Index-TTS, F5-TTS, FireRe…
621416active
R3gm/SoniTranslate
SoniTranslate is a Gradio-based web application that automatically dubs videos into other languages. It transcribes speech, translates it, …
551410active
joewongjc/type4me
Type4Me is a macOS voice input method that transcribes speech in real time (streaming, local or cloud ASR engines) and refines text with LL…
811404active
wxbool/video-srt-windows
VideoSrt is an open-source Windows GUI tool written in Go that recognizes speech in video/audio files and automatically generates SRT subti…
235032maintenance
0nutation/SpeechGPT
SpeechGPT is a series of speech large language models that can perceive and generate speech, supporting cross-modal instruction following a…
291401active
elder-plinius/V3SP3R
V3SP3R (Vesper) is an Android application that turns a Flipper Zero into an AI-controlled hardware hacking tool, driven by natural language…
481400active
zackees/transcribe-anything
A Python CLI app that transcribes local audio/video files or URLs (YouTube, Rumble, etc.) using multiple Whisper backends with automatic de…
931393active
wenet-e2e/wespeaker
WeSpeaker is a research and production-oriented toolkit for speaker embedding learning, supporting speaker verification, recognition, and d…
641392active
0xsline/OpenChatCut
OpenChatCut is an open-source, local-first desktop AI video editor with a professional multi-track timeline that AI agents (built-in, Codex…
801388active
callstackincubator/ai
A collection of on-device AI primitives for React Native offering Vercel AI SDK-compatible APIs for running LLMs, embeddings, transcription…
761387active
haoheliu/voicefixer
VoiceFixer is a Python library and CLI tool for general speech restoration, using a pretrained neural vocoder to restore degraded human spe…
261373stable
elevenyellow/handcrafted-persona-engine
Persona Engine is a Windows desktop application that drives a Live2D avatar with an AI pipeline: microphone speech recognition, an LLM guid…
731357active
k2-fsa/k2
k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)…
641352active
nyrahealth/CrisperWhisper
CrisperWhisper 2.0 is a controllable speech recognition model and Python library that transcribes audio either verbatim (including fillers,…
891349active
Softcatala/whisper-ctranslate2
A command-line transcription and translation tool compatible with OpenAI's Whisper CLI, built on CTranslate2 and faster-whisper for up to 4…
601343active
AI-FanGe/OpenAIglasses_for_Navigation
An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st…
391343active
claritylab/lucida
Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros…
324782maintenance
mbailey/voicemode
VoiceMode is a Python-based MCP server and Claude Code plugin that enables natural, hands-free voice conversations with Claude Code and oth…
801339active
mmorise/World
WORLD is a C++ library for high-quality speech analysis, manipulation, and synthesis based on a vocoder design. It estimates F0 (via DIO/Ha…
641338stable
google-gemma/gemma-translator
An on-device, fully offline voice translator application powered by Gemma 4 via LiteRT-LM, with a React web frontend optimized for small ha…
561337active
fayazara/Screendrop
Screendrop is a native macOS app for screenshots and screen recording, serving as a free, self-hostable Loom alternative. It offers annotat…
811333active
GauravSingh9356/J.A.R.V.I.S
A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,…
481332active
foobnix/LibreraReader
Librera Reader is a highly customizable e-book reader application for Android supporting PDF, EPUB, MOBI, DjVu, FB2, CBZ/CBR and many other…
974740maintenance
Nativ
Nativ is a free, MIT-licensed macOS desktop application for running open AI models locally on Apple Silicon Macs, built on MLX-VLM. It prov…
801330active
ratwithacompiler/OBS-captions-plugin
An OBS Studio plugin that generates closed captions for live streams using the Google Cloud Speech Recognition API. It leverages Twitch's b…
831324active
Robitx/gp.nvim
Gp.nvim is a Neovim plugin that brings GPT-powered chat sessions, instructable text/code operations, speech-to-text, and image generation d…
361320active
sanchit-gandhi/whisper-jax
An optimized JAX implementation of OpenAI's Whisper speech recognition model, built on Hugging Face Transformers, offering up to 70x faster…
304682maintenance
via007/bilibili-rag
A self-hosted RAG application that turns a Bilibili user's favorited videos into a searchable, conversational knowledge base. It pulls favo…
761306active
Henry-23/VideoChat
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late…
481303active
stenolabs/stenoai
Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire…
841299active
espressif/esp-box
ESP-BOX is Espressif's AIoT development framework for the ESP32-S3-BOX series of development boards, built on the ESP32-S3 Wi-Fi + Bluetoot…
611297active
ictnlp/StreamSpeech
StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p…
371287active
AudileTeam/Audile
Audile is an open-source Android app for music recognition that identifies songs playing nearby using AudD, ACRCloud, and Shazam services. …
991284active
jasperproject/jasper-client
Client code for the Jasper voice computing platform, an open source platform for building always-on, voice-controlled applications. It prov…
324518maintenance
phuc-nt/my-translator
A Tauri-based desktop app that captures system or microphone audio, transcribes it, and shows translations in a minimal real-time overlay, …
761275active
ABexit/ASR-LLM-TTS
An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) …
601271active
jordanrendric/claude-video-vision
A Claude Code plugin with an MCP server that gives Claude the ability to watch and understand videos by extracting frames via ffmpeg and tr…
581265active
UltraStar-Deluxe/USDX
UltraStar Deluxe is a free and open source karaoke singing game for PC, inspired by Sony's SingStar, where up to six (or twelve in party mo…
991260active
yzfly/douyin-mcp-server
A Python MCP server and WebUI that extracts watermark-free Douyin (TikTok China) video download links and transcribes video speech into tex…
101258active
kubeai-project/kubeai
KubeAI is a Kubernetes operator for serving machine learning models in production, supporting LLMs via vLLM and Ollama, vector embeddings, …
891256active
peteonrails/voxtype
Voxtype is a local, offline voice-to-text dictation tool for Linux with push-to-talk hotkey support across Wayland compositors like Hyprlan…
781252active
FoloUp/FoloUp
FoloUp is an open-source, self-hostable web application that conducts AI-powered voice interviews with job candidates. It generates tailore…
581250active
Blueturboguy07/cue
cue is an open-source desktop AI copilot that floats a glass panel over your screen, capturing your screen, microphone, and meeting audio t…
781247active
sh-lee-prml/HierSpeechpp
Official PyTorch implementation of HierSpeech++, a fast zero-shot speech synthesizer for text-to-speech and voice conversion based on hiera…
281238active
OStudi/short-video-generator-AI
An open-source Python tool that turns YouTube videos into ready-to-post vertical short videos by automatically detecting highlights, adding…
571224active
yeyupiaoling/Whisper-Finetune
A toolkit for fine-tuning OpenAI's Whisper speech recognition models using LoRA, supporting training with or without timestamps and even wi…
661223active
SeaDve/Mousai
Mousai is a GNOME desktop application written in Rust that identifies songs playing nearby, similar to Shazam. It listens via microphone or…
801218active
tonyqinatcmu/SlideBot-AI
SlideBot AI is an AI-powered presentation generator that turns a topic, outline, or uploaded materials (documents, spreadsheets, meeting re…
441214active
uezo/ChatdollKit
ChatdollKit is a Unity SDK that turns 3D character models into voice-enabled chatbots and virtual assistants. It integrates LLMs (ChatGPT, …
761213active
modal-labs/quillman
QuiLLMan is a voice chat application built on Kyutai's Moshi speech-to-speech language model, deployed serverlessly on Modal with a FastAPI…
681213active
hgneng/ekho
Ekho is an open-source Chinese text-to-speech engine supporting Mandarin, Cantonese, and Tibetan, part of the eGuideDog accessibility proje…
681211active
ardha27/AI-Song-Cover-RVC
A collection of Google Colab and Kaggle notebooks that form an all-in-one toolkit for creating AI song covers with RVC (Retrieval-based Voi…
681207active
aTrainTranscription/aTrain
aTrain is a desktop GUI application for offline transcription of speech recordings using Whisper-based machine learning models, with speake…
791205active
hkjarral/AVA-AI-Voice-Agent-for-Asterisk
An open-source AI voice agent that integrates with Asterisk/FreePBX phone systems via Audiosocket/RTP, built in Python with a modular pipel…
851202active
nishuzumi/gemini-teacher
A Python CLI application that acts as an English speaking practice assistant powered by Google Gemini. It listens to your speech via microp…
681200active
kizuna-ai-lab/sokuji
Sokuji is a cross-platform real-time two-way speech translation app for bilingual meetings, available as a desktop application (Windows, ma…
831199active
aydinnyunus/ai-captcha-bypass
A Python command-line tool that uses multimodal LLMs (GPT-4o, Gemini) to automatically solve various CAPTCHA types, including text, reCAPTC…
611196active
diodiogod/TTS-Audio-Suite
A ComfyUI custom node suite providing unified multi-engine Text-to-Speech, Voice Conversion, and audio editing across 19 engines like Chatt…
841185active
digimata/parrot
Parrot is an ultra-minimalist push-to-talk dictation daemon for macOS that transcribes speech entirely on-device using WhisperKit on the Ap…
731185active

← prev page 4 / 9 next →