Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: speech-recognition

801 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
jhj0517/Whisper-WebUI
A Gradio-based web interface for OpenAI's Whisper models that generates subtitles from files, YouTube videos, or microphone input. It suppo…
632862active
fqscfqj/Y2A-Auto
Y2A-Auto is a self-hosted Python application that automates re-uploading YouTube videos to AcFun and bilibili. It handles the full pipeline…
822845active
linto-ai/whisper-timestamped
A Python library extending OpenAI's Whisper models to produce accurate word-level timestamps and confidence scores during multilingual spee…
792841active
Camb-ai/MARS5-TTS
MARS5 is an open-source English text-to-speech model from CAMB.AI that uses a two-stage AR-NAR pipeline to generate expressive speech with …
142817active
rhasspy/piper
Piper is a fast, local neural text-to-speech system that runs offline on modest hardware, including Raspberry Pi devices. It offers many pr…
1011281maintenance
AutoArk/GPA
GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an…
542762active
hahahumble/speechgpt
SpeechGPT is an open-source web application that lets users have voice conversations with ChatGPT using speech recognition and speech synth…
622751active
rhasspy/rhasspy
Rhasspy is a fully offline, privacy-focused set of voice assistant services supporting many human languages. It converts spoken voice comma…
102750active
Starmel/OpenSuperWhisper
OpenSuperWhisper is a macOS dictation application that records audio and transcribes it locally using Whisper or Parakeet models. It suppor…
732721active
Vexa-ai/vexa
Vexa is an open-source (Apache-2.0) meeting transcription API that dispatches bots to join Google Meet, Microsoft Teams, and Zoom calls and…
872719active
quik-sms/quik
QUIK is an open-source SMS messenger app for Android, a revived continuation of QKSMS. It replaces the stock messaging app with features li…
872708active
dscripka/openWakeWord
openWakeWord is an open-source Python library for detecting wake words (or phrases) in streaming audio, with pre-trained models for common …
492702active
FluidInference/FluidAudio
A Swift SDK providing fully local, low-latency audio AI on Apple devices, including speech-to-text (Parakeet), speaker diarization (Pyannot…
812699active
yxlllc/DDSP-SVC
DDSP-SVC is an open-source singing voice conversion system built on Differentiable Digital Signal Processing, designed as a free AI voice c…
652656active
Const-me/Whisper
A Windows port of whisper.cpp that runs OpenAI's Whisper speech recognition model on the GPU via DirectCompute (Direct3D 11 compute shaders…
6010649maintenance
ZeframLou/call-me
A minimal Claude Code plugin that places real phone calls to notify you when an AI coding agent finishes a task, gets stuck, or needs a dec…
542638active
anliyuan/Ultralight-Digital-Human
An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It…
642627active
nvaccess/nvda
NVDA (NonVisual Desktop Access) is a free, open source screen reader for Microsoft Windows that reads on-screen text aloud via synthetic sp…
902625stable
iamsrikanthnani/pluely
Pluely is a privacy-first desktop AI assistant that runs as an invisible always-on-top overlay, providing live meeting transcription, on-sc…
792596active
liou666/polyglot
Polyglot is a cross-platform desktop (and web) application for practicing spoken language with AI conversation partners, built on ChatGPT f…
392585active
ahmedeltaher/Android-MVVM-Architecture-Android-Voice-AI-SDK
A reusable Android library (Kotlin, MVVM) that provides a full voice-driven AI conversation pipeline: microphone capture with VAD, speech-t…
702578active
zachlatta/freeflow
FreeFlow is a free, open-source macOS dictation app that transcribes speech via Groq's fast API and pastes cleaned text into any focused te…
772576active
yazinsai/OpenOats
OpenOats is a macOS meeting assistant that transcribes both sides of a call in real time using on-device speech recognition and surfaces re…
772565active
google-gemini/live-api-web-console
A React-based starter application for building real-time multimodal experiences with Google's Gemini Live API over WebSockets. It provides …
612558active
zixiiu/Digital_Life_Server
A Python server backend for a 'digital life' voice assistant that combines speech recognition, ChatGPT-based conversation, sentiment analys…
302554active
AIGC-Audio/AudioGPT
AudioGPT is a Python framework that wraps multiple audio foundation models (for speech, singing, sound, and talking-head tasks) behind a GP…
3010167maintenance
Intent-Lab/VisionClaw
VisionClaw is a real-time AI assistant app for Meta Ray-Ban smart glasses that streams camera frames and microphone audio to the Gemini Liv…
592529active
microsoft/foundry-local
Foundry Local is Microsoft's end-to-end local AI runtime and SDK suite (C#, JavaScript, Python, Rust) for running optimized models entirely…
852525active
janhq/ichigo
Ichigo is a Python speech package for developers offering local realtime voice AI capabilities, including a compact 22M-parameter speech to…
502492active
KieronQuinn/AmbientMusicMod
Ambient Music Mod is an Android app that ports the Pixel-exclusive Now Playing feature to other Android devices, using Shizuku or Sui for e…
232492active
Live-GalGame/LiveGalGame
LiveGalGame is a playful app that overlays a visual-novel (galgame) interface onto real-life conversations, providing real-time speech-to-t…
582490active
harry0703/AudioNotes
AudioNotes is a locally-run web application that transcribes audio and video files into text and organizes them into structured Markdown no…
672472active
badlogic/pi-skills
A collection of skills (SKILL.md-based tool packages) for the pi coding agent, also compatible with Claude Code, Codex CLI, Amp, and Droid.…
552434active
Alan AI SDK
Alan AI SDK is a set of client libraries for embedding Alan AI's conversational AI agents and intelligent app layer into web, iOS, Android,…
932431active
erew123/alltalk_tts
AllTalk TTS is a text-to-speech application built on the Coqui TTS engine, usable standalone or as an extension for Text-generation-webui, …
442429active
pnnbao97/VieNeu-TTS
VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi…
842427active
ace-trump-tech/MindPaw
MindPaw is an open-source desktop quadruped robot dog built on ESP8266 with roughly ¥50 in parts, featuring voice control, gesture recognit…
582427active
wan-h/awesome-digital-human-live2d
An open-source digital human application that combines Live2D avatars with LLM-powered conversation, integrating ASR, LLM, TTS, and agent o…
622413active
yetone/voice-input-src
A macOS menu-bar voice input app (Swift, macOS 14+) that lets users hold the Fn key to record and release to inject streaming-transcribed t…
502409active
earlephilhower/ESP8266Audio
An Arduino library for parsing and decoding audio formats including MOD, WAV, MP3, FLAC, MIDI, AAC, OGG/Opus, and RTTTL on ESP8266, ESP32, …
962392active
Natively-AI-assistant/natively-cluely-ai-assistant
Natively is a free, source-available desktop AI meeting assistant and interview copilot that provides real-time transcription, AI-generated…
822354active
Mentra-Community/MentraOS
MentraOS is an open-source operating system and development platform for smart glasses, providing pairing, connection management, data stre…
942326active
ggerganov/kbd-audio
A collection of command-line and GUI tools that capture and analyze microphone audio to recover keyboard keystrokes acoustically. Its Keyta…
239024maintenance
espressif/esp-adf
Espressif's official Advanced Development Framework for building audio and multimedia applications on ESP32-series SoCs. It provides pipeli…
802299active
cosin2077/easyVoice
EasyVoice is an open-source text-to-speech application that converts long texts and novels into high-quality audio with streaming playback …
492287active
yan5xu/ququ
QuQu is an open-source, free desktop voice dictation app built with Electron, designed as a privacy-first Wispr Flow alternative optimized …
312271active
QwenAudio/qwen-audio-agent
A realtime voice runtime and frontend for AI coding agents like Claude Code, Codex, and Qwen Code, letting agents talk, listen, and report …
792264active
TEN-framework/ten-vad
TEN VAD is a lightweight, low-latency voice activity detection library with an ONNX model for real-time speech detection. It supports C, Py…
502248active
floneum/kalosm
Kalosm is a Rust ecosystem of crates providing simple interfaces for running pre-trained language, audio, and image models locally or remot…
642223active
KikoPlayProject/KikoPlay
KikoPlay is a full-featured danmaku (bullet comment) video player built on Qt and libmpv, with danmu fetching from major video sites, a tre…
832219active
DigitalPhonetics/IMS-Toucan
IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive…
632207active
TransWithAI/Faster-Whisper-TransWithAI-ChickenRice
A high-performance audio/video transcription and translation application built on Faster Whisper, optimized for Japanese-to-Chinese transla…
772196active
meizhong986/WhisperJAV
WhisperJAV is a local, privacy-preserving subtitle generator for Japanese adult videos, combining Qwen3-ASR, Whisper, TEN-VAD/FireRedVAD vo…
942170active
Nutlope/notesGPT
NotesGPT is an open-source AI-powered voice note-taking web app that records voice notes, transcribes them with Whisper, and generates summ…
682153active
boson-ai/higgs-audio
Higgs Audio is a text-audio foundation model project from Boson AI providing code and weights for conversational text-to-speech with zero-s…
578329maintenance
kaixxx/noScribe
noScribe is a free, open-source desktop application that transcribes audio locally using OpenAI's Whisper (via faster-whisper) and pyannote…
782132active
kleinlee/DH_live
DH_live (mini) is an open-source 2D talking-head digital human toolkit that generates real-time lip-synced avatar video from a single refer…
672131active
QwenLM/Qwen2-Audio
Official repository for Qwen2-Audio, a 7B-parameter large audio-language model from Alibaba Cloud that accepts audio inputs and responds to…
312099active
schibsted/WAAS
Whisper as a Service (WAAS) is a self-hosted GUI and API wrapper around OpenAI's Whisper speech-to-text model, with asynchronous job queuin…
732075active
HUANGCHIHHUNGLeo/claude-real-video
A Python CLI tool and agent skill that lets LLMs like Claude actually watch videos by extracting scene-aware, deduplicated keyframes plus a…
802072active
kimjammer/Neuro
A local recreation of the Neuro-Sama AI VTuber that runs open-source LLMs on consumer hardware, combining realtime speech-to-text, text-to-…
282070active
intel/openvino-plugins-ai-audacity
A set of AI-enabled effects, generators, and analyzers for Audacity, powered by Intel OpenVINO running fully locally on CPU, GPU, or NPU. I…
732066active
Plachtaa/VALL-E-X
An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint…
107931maintenance
ricky0123/vad
A JavaScript/TypeScript library that runs Silero VAD via ONNX Runtime Web to detect voice activity directly in the browser. It provides a s…
692043active
mli/autocut
AutoCut is a Python CLI tool that automatically transcribes video audio into subtitles using Whisper, then cuts video segments based on whi…
327790maintenance
0xShug0/audio.cpp
audio.cpp is a pure C++ inference engine for audio models built on ggml, supporting TTS, STT, VAD, voice conversion, music generation, and …
802022active
juanmc2005/diart
Diart is a Python framework for building AI-powered real-time audio applications, best known for state-of-the-art streaming speaker diariza…
622022active
digitalsamba/claude-code-video-toolkit
An AI-native video production toolkit designed for Claude Code, providing skills, commands, templates, and Python tools so an AI agent can …
832008active
Niek/chatgpt-web
A single-page web interface for OpenAI-compatible chat APIs, built with Svelte, where users bring their own API key and chats are stored pr…
751996active
getopenscreen/openscreen
OpenScreen is a free, open-source desktop screen recorder and video editor for Windows, macOS, and Linux that turns raw captures into polis…
821981active
rserota/wad
WadJS is a JavaScript library built on the Web Audio API for dynamic sound synthesis and audio manipulation, supporting playback of audio f…
551981active
FireRedTeam/FireRedASR
FireRedASR is a family of open-source industrial-grade automatic speech recognition models supporting Mandarin, Chinese dialects, and Engli…
521971active
praat/praat.github.io
Praat is a desktop application for analyzing, synthesizing, and manipulating speech, widely used in phonetics research and teaching. It pro…
991965stable
kadirnar/whisper-plus
A Python library wrapping OpenAI Whisper-family models (including distil-whisper and MLX variants) for fast speech-to-text transcription wi…
601956active
mcncarl/yichen-skills
A collection of Claude Code and Codex skills for content creators, covering writing workflows, X Articles publishing, WeChat/WeCom data cap…
701943active
akdeb/ElatoAI
ElatoAI is an open-source platform for running realtime voice AI conversations on ESP32/Arduino hardware, supporting 100+ STT, LLM, and TTS…
641933active
bugbakery/audapolis
Audapolis is a free, open-source desktop editor for spoken-word audio and video that transcribes speech automatically and lets you edit med…
691890active
flybirdxx/ComfyUI-Qwen-TTS
A ComfyUI custom node plugin that wraps Alibaba's Qwen3-TTS model for speech synthesis, zero-shot voice cloning, and natural-language voice…
541876active
lanbinleo/bili2text
bili2text is a Python command-line tool that converts Bilibili videos into text transcripts from a link or BV number, handling download, au…
561873active
MontrealCorpusTools/Montreal-Forced-Aligner
Montreal Forced Aligner is a command line utility for time-aligning orthographic transcriptions and pronunciation dictionary entries to aud…
981872active
MixLabPro/comfyui-mixlab-nodes
A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec…
671863active
dindin0497/HearIt
HearIt is an Android accessibility app that captures microphone audio and transcribes it to text in real time, displaying large on-screen c…
621842active
handy-computer/transcribe.cpp
A C/C++ speech-to-text inference library built on the ggml runtime that runs 16+ ASR model families (Whisper, Parakeet, Canary, Moonshine, …
811836active
RHVoice/RHVoice
RHVoice is a free and open-source statistical parametric speech synthesizer (TTS) built on HTS technology, originally for Russian and now s…
911831active
ROCm/FastFlowLM
FastFlowLM (FLM) is an NPU-first LLM inference runtime purpose-built and deeply optimized for AMD Ryzen AI NPUs (XDNA2), offering an Ollama…
851809active
sxzxs/Real-time-translation-typing
A Windows AutoHotkey-based tool that provides real-time typing translation and real-time speech-to-text translation, with special support f…
951804active
FL33TW00D/whisper-turbo
Whisper Turbo is a fast, cross-platform, GPU-accelerated implementation of OpenAI's Whisper speech recognition model, built on the Ratchet …
191794active
abb128/LiveCaptions
LiveCaptions is a Linux desktop application that displays real-time captions for desktop or microphone audio using a local speech recogniti…
261783active
k2-fsa/sherpa-ncnn
A C++ library for real-time offline speech recognition, text-to-speech, and voice activity detection built on the ncnn inference framework …
581779active
GtOkAi/ligar-cobranca
A Node.js CLI tool that automatically places repeated voice calls to debt collection companies using the Zenvia (formerly TotalVoice) or Tw…
621777active
wwbin2017/bailing
Bailing is an open-source voice assistant application similar to GPT-4o, built with an ASR + VAD + LLM + TTS pipeline (FunASR, silero-vad, …
541757active
x007xyz/flycut-caption
FlyCut Caption is an AI-powered video subtitle editing tool built as a React component and web/desktop app, offering speech recognition wit…
781753active
TypeWhisper/typewhisper-mac
TypeWhisper is a macOS application for local, on-device speech-to-text dictation built on WhisperKit, with optional cloud transcription via…
781735active
OpenMOSS/MOSS-Transcribe-Diarize
MOSS-Transcribe-Diarize 0.9B is an open-source end-to-end audio understanding model that jointly performs multi-speaker speech transcriptio…
581695active
absadiki/subsai
Subs AI is a subtitles generation tool available as a Web-UI, CLI, and Python package, powered by OpenAI's Whisper and its variants (faster…
661682active
isair/jarvis
Jarvis is a fully local, offline AI voice assistant for your computer that supports natural conversational interaction, unlimited memory, w…
841656active
LokerL/tts-vue
TTS-Vue is a cross-platform desktop text-to-speech application built with Electron, Vue, ElementPlus, and Vite that uses Microsoft's Edge T…
586101maintenance
SevaSk/ecoute
Ecoute is a live transcription application that captures audio from both the user's microphone and speakers and displays real-time transcri…
646047maintenance
Prat011/free-cluely
An open-source Electron desktop app that acts as an invisible AI assistant, providing real-time answers and insights during meetings, inter…
531631active
difyz9/ytb2bili
ytb2bili is a self-hosted video workflow application that automates downloading videos from YouTube, generating and translating subtitles, …
791618active

← prev page 3 / 9 next →