Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: speech-recognition

801 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
goodroot/hyprwhspr
hyprwhspr is a native Linux system-wide speech-to-text dictation application supporting local models (Whisper, Parakeet, Cohere) with optio…
851182active
NVIDIA/audio-flamingo
NVIDIA's PyTorch implementation of the Audio Flamingo series of large audio-language models (AF1, AF2, AF3, and Music Flamingo) for audio u…
501182active
ddlBoJack/emotion2vec
Official PyTorch implementation of emotion2vec, a self-supervised pre-trained model for speech emotion representation. It provides code for…
271179active
rudrankriyam/Foundation-Models-Framework-Lab
A native iOS and macOS workbench app for learning, testing, and evaluating Apple's Foundation Models framework. It provides editable recipe…
821177active
nari-labs/dia2
Dia2 is a streaming dialogue text-to-speech model from Nari Labs that can generate conversational audio in realtime without needing the ful…
411172active
matthiasn/lotti
Lotti is a private, local-first logbook app for journaling, task management, time tracking, habits, and health data, with a staff of person…
981169active
wladradchenko/wunjo.wladradchenko.ru
Wunjo CE is an open-source, locally-run AI media suite for face swap, lip sync, voice cloning, object/text/background removal, restyling, a…
701169active
soniqo/speech-swift
An open-source Swift toolkit for on-device speech AI on Apple Silicon, providing ASR, TTS, speech-to-speech, VAD, and speaker diarization v…
831161active
GENEXIS-AI/chromex
Chromex is a Chrome MV3 side-panel extension that connects the browser to OpenAI's Codex CLI via a local native-messaging bridge. It lets u…
721152active
janvarev/Irene-Voice-Assistant
Irene is an offline-capable Russian voice assistant written in Python that recognizes speech via Vosk and responds using TTS engines. Its f…
651152active
TensorSpeech/TensorFlowTTS
TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me…
233995maintenance
google/lyra
Lyra is a very low-bitrate speech codec from Google that combines traditional codec techniques with generative machine learning models to c…
233973maintenance
midas-research/audino
Audino is an open-source, self-hosted web application for annotating audio, supporting transcription, labeling, and speaker-related tasks. …
521146active
facebookresearch/fairseq2
fairseq2 is a PyTorch-based sequence modeling toolkit from Meta FAIR for training custom models for content generation tasks such as langua…
891143active
andabi/deep-voice-conversion
A TensorFlow implementation of deep neural networks for voice conversion (voice style transfer) that converts a source speaker's voice into…
323938maintenance
xzf-thu/Mega-ASR
Mega-ASR is a foundation automatic speech recognition model trained on 2.6M samples spanning 7 atomic acoustic conditions and 54 compound r…
591136active
video-db/call.md
Call.md is an open-source Electron desktop app that records meetings locally, transcribes them in real time with speaker separation, and pr…
591132active
sooftware/conformer
An unofficial PyTorch implementation of the Conformer architecture (convolution-augmented Transformer) from the INTERSPEECH 2020 paper, tar…
721131active
PriesiaMioShirakana/DragonianVoice
A C++ inference library for running ONNX-based TTS, SVC (singing voice conversion), and SVS (singing voice synthesis) models, supporting ar…
401129active
PromtEngineer/Verbi
Verbi is a modular Python voice assistant application for experimenting with state-of-the-art transcription, LLM response generation, and t…
481124active
FutureUniant/Tailor
Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea…
371115active
ardha27/AI-Waifu-Vtuber
An AI waifu Vtuber application that listens to your voice via Whisper speech recognition, generates in-character replies with OpenAI, and s…
681113active
KoljaB/RealtimeVoiceChat
A Python client-server application enabling natural spoken conversations with LLMs, streaming browser audio via WebSockets through Realtime…
333832maintenance
gabber-dev/gabber
Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr…
441111active
fxy2311-youyou/expression-trainer
A local desktop app (Electron) that trains spoken expression skills by combining fully offline real-time speech recognition (Sherpa-ONNX) w…
551105active
dominostars/playtranslate
PlayTranslate is a real-time screen translation app for Android that captures game or app text via OCR and translates it, with support for …
811102active
LoredCast/filewizard
A self-hosted, browser-based web UI for converting files between many formats, running OCR on PDFs and images, transcribing audio with Whis…
461102active
vocodedev/vocode-core
Vocode is an open-source Python library for building real-time, voice-based LLM applications and agents. It orchestrates streaming transcri…
213785maintenance
Xinrea/bili-shadowreplay
A desktop and Docker-deployable tool that continuously caches live streams (Bilibili, Douyin, TikTok, Kuaishou) and lets users clip segment…
931099active
savbell/whisper-writer
WhisperWriter is a small desktop dictation app that transcribes microphone audio to text using OpenAI's Whisper model, either locally via f…
201097active
Saik0s/Whisperboard
WhisperBoard is an open-source iOS app for voice recording and transcription powered by OpenAI's Whisper model running on-device via whispe…
571093active
nobodywho-ooo/nobodywho
NobodyWho is an open-source (EUPL-1.2) on-device LLM inference engine written in Rust, built on llama.cpp, with SDKs for Kotlin, Swift, Pyt…
861086active
FujiwaraChoki/supoclip
SupoClip is an open-source, AI-powered video clipping tool that turns long videos like podcasts and streams into vertical 9:16 short clips …
621086active
olivia-ai/olivia
Olivia is an open-source chatbot written in Go that uses a neural network for natural language understanding, aiming to be a free alternati…
103717maintenance
dtsola/xiaoyaosearch
XiaoyaoSearch is a cross-platform desktop application (Electron + Python/FastAPI) that lets users find local files using AI-powered semanti…
701083active
festvox/flite
Flite (festival-lite) is a small, fast, portable run-time text-to-speech synthesis engine written entirely in ANSI C by Carnegie Mellon Uni…
231081stable
XiaomiMiMo/MiMo-Audio
Xiaomi's open-source 7B audio language model family (Base and Instruct) plus a 1.2B RVQ audio tokenizer, pretrained on 100M+ hours of audio…
571077active
Muesli-HQ/muesli
Muesli is an open-source native macOS app that combines hotkey-driven AI dictation with local meeting transcription, running speech-to-text…
821072active
ufal/whisper_streaming
A Python library that turns Whisper-like speech recognition models into a real-time streaming transcription and translation system using a …
533672maintenance
IAHispano/Applio
Applio is an open-source, MIT-licensed voice conversion suite built on RVC that lets users convert audio into other voices, train custom vo…
903644maintenance
brenpoly/be-more-agent
An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp…
641060active
Makememo/MemoAI
MemoAI is a desktop application for macOS and Windows that transcribes audio and video (YouTube links, podcasts, local files) into text and…
961059active
AnnaSuSu/TechSpar
TechSpar is an open-source AI-powered technical interview preparation application that combines targeted training, resume-based mock interv…
821055active
zai-org/GLM-TTS
GLM-TTS is a Python-based text-to-speech synthesis system built on large language models, using a two-stage LLM plus Flow model architectur…
501055active
alphacep/vosk-android-demo
A demo Android application showing offline speech recognition and speaker identification using the Vosk and Kaldi libraries. It serves as a…
481055active
Artrajz/vits-simple-api
A Python HTTP API service that exposes VITS-family text-to-speech models (VITS, Bert-VITS2, GPT-SoVITS, W2V2 emotional VITS) for inference.…
691052active
introlab/odas
ODAS (Open embeddeD Audition System) is a C library for real-time sound source localization, tracking, separation, and post-filtering using…
321048stable
Kieirra/murmure
Murmure is a privacy-first, open-source desktop speech-to-text application that transcribes voice entirely on-device using NVIDIA's Parakee…
851047active
shang-zhu/violin
Violin is an open-source video translation tool that transcribes speech, translates it into 33 languages, synthesizes a native-sounding voi…
521047active
xiaochong/hi-kid
HiKid is a free, open-source Electron desktop app that lets children in non-English-speaking countries practice English speaking and listen…
541043active
PatterAI/Patter
Patter is an open-source, MIT-licensed SDK (Python and TypeScript) that connects AI agents to real phone calls, handling telephony, speech-…
781039active
susiai/susi_chat
A chat interface for communicating with a locally hosted LLM (via llama.cpp) through a terminal console, a browser-based console, or a voic…
611036active
mybigday/llama.rn
A React Native binding of llama.cpp that enables on-device LLM inference on iOS and Android. It supports GPU/NPU acceleration (Metal, Hexag…
941026active
rayenfeng/riko_project
Project Riko is an anime-themed conversational voice assistant that combines OpenAI's GPT for dialogue, GPT-SoVITS for voice synthesis, and…
311023active
hezarai/hezar
Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta…
781013active
TensorSpeech/TensorFlowASR
TensorFlowASR is a Python library implementing automatic speech recognition architectures such as DeepSpeech2, Jasper, RNN Transducer, Cont…
661010active
HumeAI/tada
TADA is an open-source speech-language model from Hume AI that generates expressive, high-fidelity speech via text-acoustic dual alignment,…
511009active
voquill/voquill
Voquill is an open-source, cross-platform AI voice dictation app that lets users dictate into any desktop application, with AI-powered tran…
771008active
InsiderX-Pro/video-translator
An OpenClaw agent skill (written in Shell/Python) that translates and dubs videos by submitting jobs to a remote video-translation service …
541007active
AutoArk/open-audio-opd
An industrial training stack for online policy distillation (OPD) of audio models, distilling compact ASR (and planned TTS) student models …
521007active
octimot/StoryToolkitAI
StoryToolkitAI is a desktop film editing tool that transcribes, indexes, and semantically searches video footage locally, using speech reco…
651006active
Jingyi-Wu-Richael/rachel-digital-human-production
A Codex skill (plugin) that packages a repeatable workflow for producing authorized digital-human talking-head videos using MiniMax voice c…
541005active
fossasia/voxbento_translator
A Flask-based HTTP API for real-time speech-to-text transcription and translation that accepts streamed audio chunks and routes them to plu…
751002active
resemble-ai/Resemblyzer
Resemblyzer is a Python package that uses a deep learning voice encoder to convert speech audio into 256-dimensional voice embeddings. Thes…
233300maintenance
carykh/jumpcutter
A Python CLI tool that automatically edits videos by speeding up or cutting silent sections using audio volume analysis and ffmpeg. This Gi…
323149maintenance
pluja/whishper
Whishper is a self-hosted, 100% local audio transcription and subtitling suite with a web UI, powered by FasterWhisper. It transcribes audi…
613066maintenance
rishikanthc/Scriberr
Scriberr is an open-source, fully offline audio transcription application designed for self-hosters who value privacy. It transcribes audio…
702990maintenance
HaujetZhao/QuickCut
QuickCut is a lightweight, open-source video processing application built with Python and PyQt that wraps FFmpeg in a friendly GUI. It hand…
232949maintenance
pytorch/audio
TorchAudio is PyTorch's audio library providing data manipulation, transforms, and dataset loaders for audio and speech machine learning. I…
872927maintenance
tensorflow/lingvo
Lingvo is a TensorFlow-based framework for building neural networks, particularly sequence models, with a focus on speech recognition, mach…
722864maintenance
readbeyond/aeneas
aeneas is a Python/C library and set of CLI tools that automatically computes forced alignments, generating a synchronization map between a…
652863maintenance
evancohen/smart-mirror
A DIY voice-controlled smart mirror application that displays information and controls IoT smart devices, typically running on a Raspberry …
322820maintenance
marytts/marytts
MaryTTS is an open-source, multilingual text-to-speech synthesis platform written in pure Java, operating as a client-server system. It sup…
242583maintenance
zzmp/juliusjs
JuliusJS is a JavaScript port of the Julius speech recognition engine that runs entirely in the browser via a Web Worker. It transcribes us…
232565maintenance
s3prl/s3prl
S3PRL is a PyTorch toolkit for self-supervised speech pre-training and representation learning, bundling many upstream models like wav2vec …
552561maintenance
wiseman/py-webrtcvad
A Python wrapper around Google's WebRTC Voice Activity Detector, classifying short frames of 16-bit mono PCM audio as speech or non-speech.…
322496maintenance
CjangCjengh/MoeGoe
MoeGoe is an executable command-line tool for running inference with VITS text-to-speech models, supporting TTS, voice conversion, HuBERT-V…
232423maintenance
jameslyons/python_speech_features
A Python library for extracting common speech features used in automatic speech recognition, including MFCCs, filterbank energies, log filt…
232423maintenance
mravanelli/pytorch-kaldi
PyTorch-Kaldi is a toolkit for developing state-of-the-art DNN/HMM hybrid speech recognition systems, combining PyTorch-managed neural netw…
322404maintenance
cogentapps/chat-with-gpt
An open-source, self-hostable ChatGPT web app built with TypeScript and React that adds voice features via ElevenLabs text-to-speech and Op…
302352maintenance
jianfch/stable-ts
A Python library that modifies OpenAI's Whisper to produce more reliable timestamps, adding transcription, forced alignment, and audio inde…
102281maintenance
m1guelpf/auto-subtitle
A Python CLI tool that uses OpenAI's Whisper and ffmpeg to automatically generate subtitles for videos and burn them into the output file. …
322275maintenance
jarikomppa/soloud
SoLoud is a free, portable C/C++ audio engine for games with a simple 'fire and forget' API for playing sounds. It supports WAV, Ogg Vorbis…
232166maintenance
SeanNaren/deepspeech.pytorch
A PyTorch implementation of the DeepSpeech2 speech recognition model, built on PyTorch Lightning, supporting training, testing, and inferen…
232136maintenance
nobody132/masr
MASR is an end-to-end Mandarin Chinese automatic speech recognition project built on a gated convolutional neural network (similar to Wav2L…
231968maintenance
QwenLM/Qwen-Audio
Official repository for Qwen-Audio, Alibaba Cloud's large audio-language model with pretrained and chat variants. It provides model weights…
271945maintenance
julius-speech/julius
Julius is an open-source large vocabulary continuous speech recognition (LVCSR) decoder written in C, based on word N-gram language models …
351933maintenance
astorfi/lip-reading-deeplearning
A TensorFlow implementation of coupled 3D convolutional neural networks for cross audio-visual matching recognition, accompanying an IEEE A…
231904maintenance
kalliope-project/kalliope
Kalliope is a modular, always-on voice-controlled personal assistant framework written in Python, designed to run on Linux systems includin…
231772maintenance
strob/gentle
Gentle is a robust yet lenient forced aligner built on Kaldi that aligns audio speech with a known text transcript. It can be used as a Mac…
651705maintenance
bjoernkarmann/project_alias
Project Alias is an open-source Raspberry Pi-based device that acts as a 'parasite' on smart home assistants, letting users train custom wa…
321700maintenance
google/aiyprojects-raspbian
Google's Python API libraries, samples, and Raspbian system images for the AIY Projects Voice Kit and Vision Kit on Raspberry Pi. It provid…
101663maintenance
vanshg/MacAssistant
MacAssistant is a macOS application that integrates the Google Assistant into the Mac menu bar using the Google Assistant SDK. It is writte…
231603maintenance
google/uis-rnn
A Python library implementing the Unbounded Interleaved-State Recurrent Neural Network (UIS-RNN) algorithm for segmenting and clustering se…
101588maintenance
yakGPT/yakGPT
YakGPT is a locally running, browser-based ChatGPT UI that connects directly to the OpenAI API with your own key. It adds hands-free voice …
301586maintenance
fossasia/susi_chromebot
A Chrome browser extension that provides access to the SUSI.AI assistant without leaving the current tab. It opens via toolbar icon or Alt+…
101546maintenance
beyondcode/writeout.ai
A self-hosted Laravel web application that transcribes uploaded audio files using OpenAI's Whisper API and translates the transcripts via t…
711525maintenance
Tameyer41/liftoff
Liftoff is a web application that simulates technical mock interviews and provides AI-powered feedback using OpenAI Whisper for speech tran…
291523maintenance
syl22-00/pocketsphinx.js
PocketSphinx.js is a speech recognition library that runs entirely in the web browser, built by compiling the PocketSphinx C recognizer to …
321508maintenance
google/live-transcribe-speech-engine
Android client libraries from Google's Live Transcribe app for streaming real-time speech recognition via the Google Cloud Speech API. It p…
101499maintenance

← prev page 5 / 9 next →