Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: speech-processing

552 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Picovoice/porcupine
Porcupine is a highly-accurate, lightweight on-device wake word detection engine powered by deep neural networks. It enables always-listeni…
734922stable
buxuku/SmartSub
SmartSub (妙幕) is a free, open-source cross-platform desktop application that provides an end-to-end subtitle and dubbing pipeline: speech-t…
894778active
jianchang512/stt
An offline local speech-to-text tool based on faster-whisper models that transcribes audio and video files into JSON, SRT subtitles, or pla…
454766active
MoonshotAI/Kimi-Audio
Kimi-Audio is an open-source audio foundation model (7B parameters) that unifies audio understanding, generation, and speech conversation i…
314731active
WhisperSpeech/WhisperSpeech
WhisperSpeech is an open-source text-to-speech system built by inverting OpenAI's Whisper model, aiming to be 'Stable Diffusion for speech'…
564639active
Rikorose/DeepFilterNet
DeepFilterNet is a low-complexity speech enhancement framework that performs real-time noise suppression on full-band 48kHz audio using dee…
234632active
gradio-app/fastrtc
FastRTC is a Python library that turns any Python function into a real-time audio and video stream over WebRTC or WebSockets. It includes b…
584622active
fixie-ai/ultravox
Ultravox is a fast multimodal LLM that understands human speech directly without a separate ASR stage, by projecting audio into the LLM's e…
444555active
jing332/tts-server-android
An Android system-wide text-to-speech (TTS) app that provides a TTS engine service with built-in Microsoft demo API support, custom HTTP TT…
404483active
modelscope/ClearerVoice-Studio
ClearerVoice-Studio is an open-source, AI-powered speech processing toolkit from ModelScope/Alibaba offering state-of-the-art pretrained mo…
384445active
cmusphinx/pocketsphinx
PocketSphinx is Carnegie Mellon's lightweight, speaker-independent continuous speech recognition engine, available as a C library, a comman…
844335active
leetcode-mafia/cheetah
Cheetah is a macOS app that provides real-time AI coaching for software engineering interviews. It transcribes conversation audio locally w…
214257active
collabora/WhisperLive
WhisperLive is a nearly-live speech-to-text application built on OpenAI's Whisper, serving real-time transcription over WebSocket and REST …
864241active
huggingface/distil-whisper
Distil-Whisper is a distilled version of OpenAI's Whisper model for English speech recognition, offering 6x faster inference, 49% fewer par…
274112active
tmoroney/auto-subs
AutoSubs is a local-first desktop application that generates AI subtitles on-device using Whisper, Moonshine, and Parakeet models, with spe…
944091active
QwenLM/Qwen2.5-Omni
Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre…
314074active
MOSS-TTS
MOSS-TTS-Nano is an open-source 0.1B-parameter multilingual speech generation (TTS) model from MOSI.AI and the OpenMOSS team, designed for …
574031active
KoljaB/RealtimeTTS
RealtimeTTS is a Python library that converts text, generators, and LLM token streams into speech audio with low latency. It supports multi…
944016active
QwenLM/Qwen3-Omni
Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an…
523980active
digimata/quill
Quill is an ultra-minimalist, fully local macOS meeting recorder and transcriber that runs as a single Swift menu-bar binary. It records mi…
553874active
openai/openai-agents-js
A lightweight, provider-agnostic TypeScript/JavaScript framework from OpenAI for building multi-agent LLM workflows, including text agents,…
813714active
murtaza-nasir/speakr
Speakr is a self-hosted web application that transcribes audio recordings and turns them into organized, searchable notes. It supports mult…
843676active
kaldi-asr/kaldi
Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea…
5215469maintenance
speaches-ai/speaches
Speaches is an OpenAI API-compatible self-hosted server for speech-to-text (via faster-whisper), translation, and text-to-speech (via Kokor…
783621active
facebookresearch/sam-audio
SAM-Audio is Meta's foundation model for isolating any sound in audio using text, visual, or temporal prompts. This repository provides inf…
553612active
SakiRinn/LiveCaptions-Translator
A lightweight Windows application that combines the built-in Windows 11 LiveCaptions speech-to-text feature with translation APIs (includin…
773585active
Soul-AILab/SoulX-Podcast
SoulX-Podcast is the official inference codebase for a text-to-speech model that generates long-form, multi-turn, multi-speaker podcast-sty…
423535active
neonbjb/tortoise-tts
Tortoise TTS is a multi-voice text-to-speech library built on PyTorch that prioritizes highly realistic prosody and intonation. It combines…
3214870maintenance
common-voice/common-voice
Mozilla Common Voice is a web platform for crowdsourcing voice donations to build public-domain speech datasets for training voice recognit…
953484active
huangjunsen0406/py-xiaozhi
py-xiaozhi is an open-source, cross-platform multimodal AI voice assistant client written in Python, compatible with the xiaozhi-esp32 ecos…
853452active
Kedreamix/Linly-Talker
Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, …
483436active
Qwen3-ASR
Qwen3-ASR is a family of open-source speech recognition models from Alibaba's Qwen team, supporting ASR and language identification across …
553423active
WEIFENG2333/AsrTools
AsrTools is a Python desktop application with a PyQt5-based GUI that converts audio and video files to text using online ASR engines, witho…
483423active
libAudioFlux/audioFlux
audioFlux is a C-based library with Python bindings for audio and music analysis and feature extraction. It supports dozens of time-frequen…
533351active
xenova/whisper-web
A browser-based speech recognition app that runs OpenAI's Whisper models entirely client-side using Transformers.js. It transcribes audio w…
303338active
ahmetoner/whisper-asr-webservice
A Dockerized REST webservice that wraps OpenAI Whisper (plus Faster Whisper and WhisperX engines) for automatic speech recognition. It expo…
933326active
Open-Less/openless
OpenLess is an open-source cross-platform voice input application that lets users hold a hotkey, speak, and have AI-polished text inserted …
813321active
xiph/opus
libopus is the reference implementation of the Opus audio codec, an IETF-standardized (RFC 6716) royalty-free codec for interactive speech …
793293stable
rsxdalv/TTS-WebUI
TTS WebUI is a free, open-source web interface combining Gradio and React that unifies 30+ AI models for text-to-speech, voice conversion, …
893244active
echo-loop/Echo-Loop
Echo Loop is an open-source Flutter-based English listening and speaking training app that guides learners through a structured listen-to-s…
803242active
VOICEVOX
VOICEVOX is a free, mid-quality text-to-speech (and singing synthesis) software whose editor is built with Electron, TypeScript, and Vue. I…
933231active
MisoLabsAI/MisoTTS
Miso TTS 8B is an open-source text-to-speech model based on an RVQ Transformer architecture with a Llama 3.2-style 8B backbone, designed fo…
523224active
zai-org/GLM-4-Voice
GLM-4-Voice is an end-to-end bilingual (Chinese/English) speech dialogue model from Zhipu AI, built on GLM-4-9B with a speech tokenizer and…
223222active
wendy7756/AI-Video-Transcriber
An open-source AI tool that transcribes, summarizes, and archives videos and podcasts from 30+ platforms (YouTube, TikTok, Bilibili, etc.) …
623215active
Purfview/whisper-standalone-win
Standalone Windows/Linux/macOS executables of OpenAI's Whisper and Faster-Whisper that transcribe audio and video to text and subtitles wit…
503163active
BayLing-Models/BayLing-Speech
LLaMA-Omni is an end-to-end speech interaction model built on Llama-3.1-8B-Instruct that generates simultaneous text and speech responses f…
323146active
matthartman/ghost-pepper
A free, open-source macOS menu bar app providing fully on-device speech-to-text dictation and meeting transcription using local Whisper, Pa…
783140active
modelscope/3D-Speaker
3D-Speaker is an open-source Python toolkit for single- and multi-modal speaker verification, speaker recognition, and speaker diarization,…
563121active
HeyWillow/willow
Willow is an open source, self-hosted voice assistant platform for ESP32-S3-BOX hardware, designed as a privacy-focused alternative to Amaz…
863097active
elevenlabs/elevenlabs-python
The official Python SDK for the ElevenLabs API, providing programmatic access to text-to-speech, speech-to-text, voice cloning, dubbing, mu…
933078active
kyutai-labs/delayed-streams-modeling
Kyutai's repository of Speech-To-Text and Text-To-Speech models built on the Delayed Streams Modeling framework, with implementations in Py…
473017active
CheshireCC/faster-whisper-GUI
A desktop GUI application built with PySide6 for running faster-whisper and whisperX speech-to-text transcription. It lets users transcribe…
712991active
KevinWang676/Bark-Voice-Cloning
A one-click hub of Gradio Web UIs and Colab notebooks for open-source voice cloning, TTS, and voice conversion models including Bark, GPT-S…
672947active
davabase/whisper_real_time
A Python demo application that performs real-time speech-to-text transcription using OpenAI's Whisper model. It records audio continuously …
392941active
facebookresearch/omnilingual-asr
An open-source multilingual speech recognition library from Meta AI supporting over 1,600 languages, including hundreds never previously co…
522898active
kitlangton/Hex
Hex is a macOS application that converts your voice to text: press-and-hold a global hotkey to record, and it transcribes on-device and pas…
852890active
openai/openai-fm
OpenAI.fm is an interactive web demo showcasing OpenAI's text-to-speech models, built with Next.js and the OpenAI Speech API. It lets users…
512887active
jhj0517/Whisper-WebUI
A Gradio-based web interface for OpenAI's Whisper models that generates subtitles from files, YouTube videos, or microphone input. It suppo…
632862active
linto-ai/whisper-timestamped
A Python library extending OpenAI's Whisper models to produce accurate word-level timestamps and confidence scores during multilingual spee…
792841active
Camb-ai/MARS5-TTS
MARS5 is an open-source English text-to-speech model from CAMB.AI that uses a two-stage AR-NAR pipeline to generate expressive speech with …
142817active
rhasspy/piper
Piper is a fast, local neural text-to-speech system that runs offline on modest hardware, including Raspberry Pi devices. It offers many pr…
1011281maintenance
AutoArk/GPA
GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an…
542762active
hahahumble/speechgpt
SpeechGPT is an open-source web application that lets users have voice conversations with ChatGPT using speech recognition and speech synth…
622751active
rhasspy/rhasspy
Rhasspy is a fully offline, privacy-focused set of voice assistant services supporting many human languages. It converts spoken voice comma…
102750active
Starmel/OpenSuperWhisper
OpenSuperWhisper is a macOS dictation application that records audio and transcribes it locally using Whisper or Parakeet models. It suppor…
732721active
Vexa-ai/vexa
Vexa is an open-source (Apache-2.0) meeting transcription API that dispatches bots to join Google Meet, Microsoft Teams, and Zoom calls and…
872719active
dscripka/openWakeWord
openWakeWord is an open-source Python library for detecting wake words (or phrases) in streaming audio, with pre-trained models for common …
492702active
FluidInference/FluidAudio
A Swift SDK providing fully local, low-latency audio AI on Apple devices, including speech-to-text (Parakeet), speaker diarization (Pyannot…
812699active
thewh1teagle/kokoro-onnx
A Python library that runs the Kokoro text-to-speech model via ONNX Runtime, supporting CPU and GPU inference. It provides multi-language T…
672680active
Const-me/Whisper
A Windows port of whisper.cpp that runs OpenAI's Whisper speech recognition model on the GPU via DirectCompute (Direct3D 11 compute shaders…
6010649maintenance
ZeframLou/call-me
A minimal Claude Code plugin that places real phone calls to notify you when an AI coding agent finishes a task, gets stuck, or needs a dec…
542638active
anliyuan/Ultralight-Digital-Human
An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It…
642627active
pndurette/gTTS
gTTS is a Python library and CLI tool that interfaces with Google Translate's text-to-speech API to generate spoken MP3 audio from text. It…
572627stable
6drf21e/ChatTTS_colab
A one-click deployment wrapper around ChatTTS providing a Gradio web UI for text-to-speech, runnable in Google Colab or via an offline Wind…
532592active
asteroid-team/asteroid
Asteroid is a PyTorch-based audio source separation toolkit for researchers, providing modular building blocks (filterbanks, encoders, mask…
602584active
ahmedeltaher/Android-MVVM-Architecture-Android-Voice-AI-SDK
A reusable Android library (Kotlin, MVVM) that provides a full voice-driven AI conversation pipeline: microphone capture with VAD, speech-t…
702578active
zachlatta/freeflow
FreeFlow is a free, open-source macOS dictation app that transcribes speech via Groq's fast API and pastes cleaned text into any focused te…
772576active
zixiiu/Digital_Life_Server
A Python server backend for a 'digital life' voice assistant that combines speech recognition, ChatGPT-based conversation, sentiment analys…
302554active
AIGC-Audio/AudioGPT
AudioGPT is a Python framework that wraps multiple audio foundation models (for speech, singing, sound, and talking-head tasks) behind a GP…
3010167maintenance
mozilla/TTS
A deep learning library for advanced text-to-speech generation, built on PyTorch with models like Tacotron2, Glow-TTS, and various vocoders…
2310167maintenance
nateshmbhat/pyttsx3
pyttsx3 is an offline text-to-speech synthesis library for Python 3 that wraps system TTS engines like Sapi5, NSSpeechSynthesizer, and espe…
672529active
microsoft/foundry-local
Foundry Local is Microsoft's end-to-end local AI runtime and SDK suite (C#, JavaScript, Python, Rust) for running optimized models entirely…
852525active
janhq/ichigo
Ichigo is a Python speech package for developers offering local realtime voice AI capabilities, including a compact 22M-parameter speech to…
502492active
Live-GalGame/LiveGalGame
LiveGalGame is a playful app that overlays a visual-novel (galgame) interface onto real-life conversations, providing real-time speech-to-t…
582490active
harry0703/AudioNotes
AudioNotes is a locally-run web application that transcribes audio and video files into text and organizes them into structured Markdown no…
672472active
erew123/alltalk_tts
AllTalk TTS is a text-to-speech application built on the Coqui TTS engine, usable standalone or as an extension for Text-generation-webui, …
442429active
pnnbao97/VieNeu-TTS
VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi…
842427active
yetone/voice-input-src
A macOS menu-bar voice input app (Swift, macOS 14+) that lets users hold the Fn key to record and release to inject streaming-transcribed t…
502409active
resemble-ai/resemble-enhance
Resemble Enhance is an AI-powered Python tool that improves speech quality through denoising and enhancement, using a denoiser module and a…
172397active
jik876/hifi-gan
The official PyTorch implementation of HiFi-GAN, a generative adversarial network that converts mel-spectrograms into high-fidelity 22.05 k…
322367stable
iver56/audiomentations
Audiomentations is a Python library for audio data augmentation with an API inspired by albumentations. It provides fast CPU-based waveform…
672314stable
cosin2077/easyVoice
EasyVoice is an open-source text-to-speech application that converts long texts and novels into high-quality audio with streaming playback …
492287active
yan5xu/ququ
QuQu is an open-source, free desktop voice dictation app built with Electron, designed as a privacy-first Wispr Flow alternative optimized …
312271active
QwenAudio/qwen-audio-agent
A realtime voice runtime and frontend for AI coding agents like Claude Code, Codex, and Qwen Code, letting agents talk, listen, and report …
792264active
fishaudio/Bert-VITS2
Bert-VITS2 is a text-to-speech model implementation combining the VITS2 architecture with multilingual BERT embeddings, written in Python. …
638796maintenance
TEN-framework/ten-vad
TEN VAD is a lightweight, low-latency voice activity detection library with an ONNX model for real-time speech detection. It supports C, Py…
502248active
lifeiteng/vall-e
An unofficial PyTorch implementation of VALL-E, a zero-shot text-to-speech model that treats TTS as a conditional language modeling task ov…
402215active
DigitalPhonetics/IMS-Toucan
IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive…
632207active
TransWithAI/Faster-Whisper-TransWithAI-ChickenRice
A high-performance audio/video transcription and translation application built on Faster Whisper, optimized for Japanese-to-Chinese transla…
772196active
meizhong986/WhisperJAV
WhisperJAV is a local, privacy-preserving subtitle generator for Japanese adult videos, combining Qwen3-ASR, Whisper, TEN-VAD/FireRedVAD vo…
942170active

← prev page 2 / 6 next →