Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: speech-recognition

801 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
microsoft/markitdown
MarkItDown is a lightweight Python utility from Microsoft that converts many file formats (PDF, Office documents, images, audio, HTML, EPub…
83176487active
openai/whisper
OpenAI's Whisper is a general-purpose speech recognition model and Python library built on a Transformer sequence-to-sequence architecture.…
68107981stable
RVC-Boss/GPT-SoVITS
GPT-SoVITS is a Python-based few-shot voice cloning and text-to-speech system with an integrated WebUI. It supports zero-shot TTS from a 5-…
6861255active
microsoft/VibeVoice
VibeVoice is Microsoft's open-source frontier voice AI framework combining a next-token diffusion text-to-speech model for expressive, long…
6053225active
ggml-org/whisper.cpp
A high-performance C/C++ port of OpenAI's Whisper automatic speech recognition model with no external dependencies. It supports CPU, GPU (M…
9953203stable
jamiepine/voicebox
Voicebox is a free, open-source, local-first AI voice studio desktop app that combines voice cloning, text-to-speech across 7 engines, and …
7551551active
moeru-ai/airi
Project AIRI is an open-source, self-hosted AI companion (waifu/virtual character) platform inspired by Neuro-sama, supporting Live2D and V…
8648451active
2noise/ChatTTS
ChatTTS is a generative text-to-speech model optimized for conversational dialogue scenarios such as LLM assistants, supporting English and…
6939797active
RVC-Project/Retrieval-based-Voice-Conversion-WebUI
A WebUI-based framework for training and running retrieval-based voice conversion (RVC) models, letting users clone a voice timbre from as …
8237845active
myshell-ai/OpenVoice
OpenVoice is a Python library and audio foundation model for instant voice cloning, requiring only a short reference audio clip to replicat…
3437312active
ZuodaoTech/everyone-can-use-english
Everyone Can Use English is an open-source project combining a well-known English learning book with the Enjoy app, an AI-powered language …
7436774active
OpenBMB/VoxCPM
VoxCPM is a tokenizer-free text-to-speech system built on a diffusion autoregressive architecture that generates continuous speech represen…
7936145active
fishaudio/fish-speech
Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 l…
7432413active
cjpais/Handy
Handy is a free, open-source, cross-platform desktop application for offline speech-to-text dictation. Press a configurable shortcut, speak…
8330402active
Zackriya-Solutions/meetily
Meetily is a privacy-first, open-source AI meeting assistant that records, transcribes, and summarizes meetings entirely locally using Para…
7929940active
resemble-ai/chatterbox
Chatterbox is a family of state-of-the-art open-source text-to-speech models by Resemble AI, including multilingual and low-latency Turbo v…
5226163active
mozilla-ai/llamafile
llamafile is a Mozilla project that packages llama.cpp (and whisper.cpp for whisperfile) with Cosmopolitan Libc into single-file executable…
9425694active
SYSTRAN/faster-whisper
A fast reimplementation of OpenAI's Whisper speech-to-text model built on the CTranslate2 inference engine, offering up to 4x speedup and l…
5725104active
m-bain/whisperX
WhisperX is a Python library providing fast automatic speech recognition with word-level timestamps and speaker diarization. It combines Wh…
9123765active
index-tts/index-tts
IndexTTS is an industrial-level zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest IndexTTS…
7823511active
QwenAudio/CosyVoice
CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-…
6122925active
browser-use/video-use
An open-source agent skill that lets coding agents like Claude Code edit raw video footage into finished MP4s via natural-language chat. It…
5921397active
chidiwilliams/buzz
Buzz is a cross-platform desktop application that transcribes and translates audio and video offline using OpenAI's Whisper models. It offe…
9521147active
w-okada/voice-changer
VCClient is a real-time AI voice changer application that converts microphone input using models like RVC and Beatrice v2. It ships as preb…
6520864active
modelscope/FunASR
FunASR is an industrial end-to-end speech recognition toolkit built on PyTorch, offering ASR, VAD, punctuation restoration, speaker diariza…
9920036active
jianchang512/pyvideotrans
pyVideoTrans is an open-source desktop application (with WebUI and CLI modes) that translates videos from one language to another. It provi…
9018812active
NVIDIA-NeMo/Speech
NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, …
9818337active
Huanshere/VideoLingo
VideoLingo is an all-in-one AI video translation, localization, and dubbing tool that produces Netflix-quality single-line subtitles with w…
8018266active
leon-ai/leon
Leon is an open-source personal AI assistant that runs on your own server, combining tools, memory, and agentic execution with speech-to-te…
8617462active
bradautomates/claude-video
A Claude Code plugin and Agent Skill that lets AI coding agents watch videos. It downloads videos via yt-dlp, extracts scene-aware frames w…
7416260active
WEIFENG2333/VideoCaptioner
VideoCaptioner is an open-source AI-powered video subtitling tool that handles the full pipeline: speech recognition, LLM-based subtitle se…
8215764active
KittenML/KittenTTS
KittenTTS is an open-source, ultra-lightweight text-to-speech library built on ONNX, with models ranging from 15M to 80M parameters (25-80 …
6615403active
Vosk
Vosk is an offline open source speech recognition toolkit supporting 20+ languages with small models and a streaming API. It provides bindi…
6615077stable
duixcom/Duix-Avatar
Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea…
5714871active
SesameAILabs/csm
CSM (Conversational Speech Model) is Sesame's speech generation model that produces conversational audio from text and audio context, using…
3014720active
k2-fsa/sherpa-onnx
sherpa-onnx is an offline speech processing toolkit built on next-gen Kaldi and onnxruntime, supporting speech-to-text, text-to-speech, spe…
9514411active
SubtitleEdit/subtitleedit
Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac…
9413971active
Open-LLM-VTuber/Open-LLM-VTuber
Open-LLM-VTuber is a cross-platform AI companion application that enables real-time voice conversations with LLMs, featuring a Live2D avata…
6713473active
Omi
Omi is an open-source AI wearable and companion app ecosystem (necklace pendant, smart glasses, desktop and mobile apps) that captures conv…
8813264active
livekit/agents
LiveKit Agents is an open-source framework for building realtime, multimodal voice AI agents that run as programmable participants in LiveK…
9013176active
QwenLM/Qwen3-TTS
Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,…
4813113active
Vaibhavs10/insanely-fast-whisper
A CLI tool for transcribing audio files on-device using OpenAI's Whisper models, powered by Hugging Face Transformers, Optimum, and Flash A…
4913049active
huggingface/speech-to-speech
A low-latency, fully modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) built from open-source models, exposed via the OpenAI Realtime…
8512882active
PaddlePaddle/PaddleSpeech
PaddleSpeech is an open-source speech and audio toolkit built on the PaddlePaddle deep learning platform, covering ASR with punctuation, st…
6612670active
abus-aikorea/voice-pro
Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, …
8412647active
facebookresearch/seamless_communication
A library of foundational multilingual multimodal AI models from Meta for speech and text translation, including SeamlessM4T, SeamlessExpre…
7111842active
rany2/edge-tts
A Python module and CLI that lets you use Microsoft Edge's online neural text-to-speech service without needing Edge, Windows, or an API ke…
7911803active
speechbrain/speechbrain
SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained…
8311785active
debpalash/VoiceStudio
VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, vid…
8011737active
ankidroid/Anki-Android
AnkiDroid is the Android client for the open-source Anki spaced repetition flashcard system, letting users study and sync decks with AnkiWe…
9311628active
CorentinJ/Real-Time-Voice-Cloning
A Python implementation of the SV2TTS (Transfer Learning from Speaker Verification to Multispeaker TTS) framework that clones a voice from …
6460110maintenance
krillinai/KrillinAI
KrillinAI is an AI-powered video translation and dubbing tool covering the full pipeline: video download, speech transcription, subtitle tr…
8411272active
TEN-framework/ten-framework
TEN is an open-source framework for building real-time multimodal conversational AI agents, with a focus on low-latency voice assistants. I…
8511082active
SparkAudio/Spark-TTS
Spark-TTS is an LLM-based text-to-speech system built on Qwen2.5 that generates speech directly from single-stream decoupled speech tokens.…
2711006active
QuentinFuxa/WhisperLiveKit
WhisperLiveKit is a self-hosted, ultra-low-latency real-time speech-to-text pipeline built on state-of-the-art simultaneous speech research…
8710962active
kyutai-labs/moshi
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai, built around the Mimi streaming neural audio…
6210949active
altic-dev/FluidVoice
FluidVoice is an open-source (GPLv3) macOS dictation app that transcribes speech to text entirely on-device using models like Parakeet, Whi…
8410946active
moonshine-ai/moonshine
Moonshine Voice is an open-source on-device AI toolkit providing very low latency speech-to-text, intent recognition, and text-to-speech fo…
8910944active
pyannote/pyannote-audio
pyannote.audio is an open-source Python toolkit built on PyTorch for speaker diarization, providing neural building blocks like voice activ…
9510475active
xinnan-tech/xiaozhi-esp32-server
A self-hosted backend server for the xiaozhi-esp32 open-source smart hardware project, implementing the Xiaozhi communication protocol over…
8110433active
NVIDIA/personaplex
PersonaPlex is a real-time, full-duplex speech-to-speech conversational model from NVIDIA that supports persona control via text role promp…
4710393active
nexmoe/VidBee
VidBee is a free, open-source Electron desktop app for downloading video and audio from 1000+ sites (via yt-dlp) and importing local media.…
8410386active
niedev/RTranslator
RTranslator is a free, open-source, offline real-time translation app for Android that runs speech recognition (Whisper) and translation (M…
8510355active
RunanywhereAI/runanywhere-sdks
RunAnywhere is a set of cross-platform SDKs (Swift, Kotlin, React Native, Flutter, TypeScript, C++) over a shared C++ core for running AI m…
8510282active
KoljaB/RealtimeSTT
RealtimeSTT is a Python library for low-latency, real-time speech-to-text with voice activity detection, wake word activation, and streamin…
9410080active
snakers4/silero-vad
Silero VAD is a pre-trained, enterprise-grade Voice Activity Detector model available via PyPI, runnable with PyTorch or ONNX Runtime. It d…
8610061stable
espnet/espnet
ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance…
889941active
xorbitsai/inference
Xinference is an open-source model serving platform for deploying LLMs, embedding, speech, image, and multimodal models via a unified OpenA…
919523active
k2-fsa/OmniVoice
OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style a…
799455active
dessant/buster
Buster is a browser extension that solves reCAPTCHA challenges by completing their audio challenges using speech recognition. It works on C…
919276active
fastrepl/anarlog
Anarlog is an open-source, local-first AI meeting notepad and Granola alternative that records device audio without a bot joining the call,…
849181active
QwenAudio/SenseVoice
SenseVoice is an open-source speech foundation model (SenseVoiceSmall) providing multilingual ASR, spoken language identification, speech e…
889151active
zyronon/TypeWords
TypeWords is a free, open-source web application for learning English vocabulary and memorizing articles through typing practice. It uses t…
809007active
jianchang512/clone-voice
A voice cloning tool with a web interface built on the coqui.ai xtts_v2 model, letting users synthesize speech in any voice from text or co…
108989active
Uberi/speech_recognition
A Python library for performing speech recognition with support for multiple engines and APIs, both online (Google, Azure, Wit.ai, OpenAI W…
948987active
Foliate
Foliate is a free, open-source e-book reader application for Linux built with GTK4 and GJS, supporting EPUB, Mobipocket, Kindle, FB2, CBZ, …
648667active
jasonppy/VoiceCraft
VoiceCraft is a token infilling neural codec language model for zero-shot speech editing and text-to-speech on in-the-wild data like audiob…
638574active
hexgrad/kokoro
An inference library for Kokoro-82M, an open-weight text-to-speech model with 82 million parameters that delivers quality comparable to lar…
368570active
netease-youdao/EmotiVoice
EmotiVoice is an open-source text-to-speech engine supporting English and Chinese with over 2000 voices and prompt-controlled emotional syn…
178523active
nl8590687/ASRT_SpeechRecognition
ASRT is a deep-learning-based Chinese speech recognition (speech-to-text) system built with TensorFlow/Keras, using CNN, LSTM, attention me…
578383active
duixcom/Duix-Mobile
Duix Mobile is an open-source SDK for building real-time interactive AI avatars (digital humans) that run on-device on Android, iOS, tablet…
698198active
GetStream/Vision-Agents
An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O…
848100active
fonoster/fonoster
Fonoster is an open-source programmable telecommunications stack (a Twilio alternative) for building voice and messaging applications, with…
998080active
smacke/ffsubsync
FFsubsync is a Python command-line tool that automatically synchronizes subtitle files (e.g., SRT) with video or audio, using voice activit…
847854active
Blaizzy/mlx-audio
MLX-Audio is a Python library built on Apple's MLX framework for fast text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS…
887793active
babysor/MockingBird
MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real…
5436909maintenance
myshell-ai/MeloTTS
MeloTTS is a high-quality multi-lingual text-to-speech library supporting English (multiple accents), Spanish, French, Chinese, Japanese, a…
167608active
pickle-com/glass
Glass by Pickle is an open-source Electron desktop app that acts as a 'digital mind extension', seeing your screen and listening in real ti…
427589active
farzaa/clicky
Clicky is an open-source macOS app that acts as an AI teacher living next to your cursor — it can see your screen, talk with you via voice,…
507394active
ilyhalight/voice-over-translation
A userscript/browser extension that adds AI voice-over translation and subtitles to videos on many websites, powered by Yandex services. It…
987335active
Zyphra/Zonos
Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I…
247244active
thewh1teagle/vibe
Vibe is a cross-platform desktop app for fully offline audio and video transcription using Whisper, Nemotron, and Parakeet models. It suppo…
937206active
JefferyHcool/BiliNote
BiliNote is an open-source AI video note-taking application that turns video links from Bilibili, YouTube, Douyin, and local files into str…
867169active
wzpan/wukong-robot
wukong-robot is a modular Chinese-language voice assistant / smart speaker project in Python that combines offline wake-word detection, ASR…
327126active
ddean2009/MoneyPrinterPlus
A Python desktop application that uses AI LLMs to batch-generate short videos with one click, including automatic video mashup/remixing and…
307008active
openai/openai-realtime-agents
A Next.js TypeScript demo application showcasing advanced agentic patterns (Chat-Supervisor and Sequential Handoff) built on the OpenAI Rea…
486966active
yihong0618/xiaogpt
A Python application that connects Xiaomi AI Speakers to ChatGPT and other LLMs, letting users ask questions starting with a trigger phrase…
646901active
TalAter/annyang
annyang is a tiny (2 KB), dependency-free JavaScript library that adds speech recognition and voice commands to websites using the Web Spee…
786816stable
werman/noise-suppression-for-voice
A real-time voice noise suppression audio plugin based on Xiph's RNNoise, available in VST2, VST3, LV2, LADSPA, AU, and AUv3 formats. It su…
776806active
Dooy/chatgpt-web-midjourney-proxy
A unified web/desktop UI for ChatGPT plus AI image, music, and video generation services like Midjourney, Suno, Luma, Runway, and Flux. It …
766785active

page 1 / 9 next →