Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: speech-processing

552 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Nutlope/notesGPT
NotesGPT is an open-source AI-powered voice note-taking web app that records voice notes, transcribes them with Whisper, and generates summ…
682153active
boson-ai/higgs-audio
Higgs Audio is a text-audio foundation model project from Boson AI providing code and weights for conversational text-to-speech with zero-s…
578329maintenance
kaixxx/noScribe
noScribe is a free, open-source desktop application that transcribes audio locally using OpenAI's Whisper (via faster-whisper) and pyannote…
782132active
kleinlee/DH_live
DH_live (mini) is an open-source 2D talking-head digital human toolkit that generates real-time lip-synced avatar video from a single refer…
672131active
QwenLM/Qwen2-Audio
Official repository for Qwen2-Audio, a 7B-parameter large audio-language model from Alibaba Cloud that accepts audio inputs and responds to…
312099active
schibsted/WAAS
Whisper as a Service (WAAS) is a self-hosted GUI and API wrapper around OpenAI's Whisper speech-to-text model, with asynchronous job queuin…
732075active
kimjammer/Neuro
A local recreation of the Neuro-Sama AI VTuber that runs open-source LLMs on consumer hardware, combining realtime speech-to-text, text-to-…
282070active
travisvn/openai-edge-tts
A self-hosted Python service that emulates the OpenAI text-to-speech API endpoint (/v1/audio/speech) using Microsoft Edge's free online TTS…
262066active
Plachtaa/VALL-E-X
An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint…
107931maintenance
jaywalnut310/vits
VITS is the official PyTorch implementation of an end-to-end text-to-speech model based on a conditional variational autoencoder with adver…
327889maintenance
ricky0123/vad
A JavaScript/TypeScript library that runs Silero VAD via ONNX Runtime Web to detect voice activity directly in the browser. It provides a s…
692043active
mli/autocut
AutoCut is a Python CLI tool that automatically transcribes video audio into subtitles using Whisper, then cuts video segments based on whi…
327790maintenance
0xShug0/audio.cpp
audio.cpp is a pure C++ inference engine for audio models built on ggml, supporting TTS, STT, VAD, voice conversion, music generation, and …
802022active
juanmc2005/diart
Diart is a Python framework for building AI-powered real-time audio applications, best known for state-of-the-art streaming speaker diariza…
622022active
FireRedTeam/FireRedASR
FireRedASR is a family of open-source industrial-grade automatic speech recognition models supporting Mandarin, Chinese dialects, and Engli…
521971active
praat/praat.github.io
Praat is a desktop application for analyzing, synthesizing, and manipulating speech, widely used in phonetics research and teaching. It pro…
991965stable
kadirnar/whisper-plus
A Python library wrapping OpenAI Whisper-family models (including distil-whisper and MLX variants) for fast speech-to-text transcription wi…
601956active
akdeb/ElatoAI
ElatoAI is an open-source platform for running realtime voice AI conversations on ESP32/Arduino hardware, supporting 100+ STT, LLM, and TTS…
641933active
LCAV/pyroomacoustics
Pyroomacoustics is a Python package for audio signal processing in indoor scenarios, combining a fast C++ room acoustics simulator (image s…
901928stable
bugbakery/audapolis
Audapolis is a free, open-source desktop editor for spoken-word audio and video that transcribes speech automatically and lets you edit med…
691890active
flybirdxx/ComfyUI-Qwen-TTS
A ComfyUI custom node plugin that wraps Alibaba's Qwen3-TTS model for speech synthesis, zero-shot voice cloning, and natural-language voice…
541876active
timsainb/noisereduce
A Python library for noise reduction in time-domain signals using spectral gating, supporting stationary and non-stationary noise reduction…
391874active
lanbinleo/bili2text
bili2text is a Python command-line tool that converts Bilibili videos into text transcripts from a link or BV number, handling download, au…
561873active
MontrealCorpusTools/Montreal-Forced-Aligner
Montreal Forced Aligner is a command line utility for time-aligning orthographic transcriptions and pronunciation dictionary entries to aud…
981872active
dindin0497/HearIt
HearIt is an Android accessibility app that captures microphone audio and transcribes it to text in real time, displaying large on-screen c…
621842active
handy-computer/transcribe.cpp
A C/C++ speech-to-text inference library built on the ggml runtime that runs 16+ ASR model families (Whisper, Parakeet, Canary, Moonshine, …
811836active
RHVoice/RHVoice
RHVoice is a free and open-source statistical parametric speech synthesizer (TTS) built on HTS technology, originally for Russian and now s…
911831active
nazdridoy/kokoro-tts
A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w…
881816active
FL33TW00D/whisper-turbo
Whisper Turbo is a fast, cross-platform, GPU-accelerated implementation of OpenAI's Whisper speech recognition model, built on the Ratchet …
191794active
abb128/LiveCaptions
LiveCaptions is a Linux desktop application that displays real-time captions for desktop or microphone audio using a local speech recogniti…
261783active
k2-fsa/sherpa-ncnn
A C++ library for real-time offline speech recognition, text-to-speech, and voice activity detection built on the ncnn inference framework …
581779active
wwbin2017/bailing
Bailing is an open-source voice assistant application similar to GPT-4o, built with an ASR + VAD + LLM + TTS pipeline (FunASR, silero-vad, …
541757active
High-Logic/Genie-TTS
GENIE is a lightweight Python inference engine for the open-source GPT-SoVITS text-to-speech project, optimized for fast CPU-based speech s…
611754active
x007xyz/flycut-caption
FlyCut Caption is an AI-powered video subtitle editing tool built as a React component and web/desktop app, offering speech recognition wit…
781753active
TypeWhisper/typewhisper-mac
TypeWhisper is a macOS application for local, on-device speech-to-text dictation built on WhisperKit, with optional cloud transcription via…
781735active
AutoArk/EVA-OS
EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining …
681712active
OpenMOSS/MOSS-Transcribe-Diarize
MOSS-Transcribe-Diarize 0.9B is an open-source end-to-end audio understanding model that jointly performs multi-speaker speech transcriptio…
581695active
absadiki/subsai
Subs AI is a subtitles generation tool available as a Web-UI, CLI, and Python package, powered by OpenAI's Whisper and its variants (faster…
661682active
isair/jarvis
Jarvis is a fully local, offline AI voice assistant for your computer that supports natural conversational interaction, unlimited memory, w…
841656active
LokerL/tts-vue
TTS-Vue is a cross-platform desktop text-to-speech application built with Electron, Vue, ElementPlus, and Vite that uses Microsoft's Edge T…
586101maintenance
SevaSk/ecoute
Ecoute is a live transcription application that captures audio from both the user's microphone and speakers and displays real-time transcri…
646047maintenance
Lynpoint/CyberVerse
CyberVerse is an open-source, self-hosted framework for building real-time, voice-first AI agents with optional digital-human video (talkin…
631612active
mkiol/dsnote
Speech Note is a Linux desktop and Sailfish OS application for note taking, reading, and translating text using offline Speech to Text, Tex…
961607active
SAGAR-TAMANG/friday-tony-stark-demo
F.R.I.D.A.Y. is a Tony Stark-inspired voice AI assistant consisting of a FastMCP server exposing tools (web search, news, system info) and …
551595active
finnvoor/yap
yap is a Swift CLI for on-device speech transcription of audio and video files using Apple's Speech.framework on macOS 26. It also supports…
811586active
royshil/obs-localvocal
LocalVocal is an OBS Studio plugin that performs real-time, fully local speech recognition and translation using Whisper models running on …
811585active
bootphon/phonemizer
A Python library and command-line tool that converts words and texts into phonemes across many languages. It wraps multiple backends (espea…
851567stable
byjlw/video-analyzer
A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs …
601559active
Enemyx-net/VibeVoice-ComfyUI
A ComfyUI custom node integration for Microsoft's VibeVoice text-to-speech model, providing single and multi-speaker voice synthesis with v…
531549active
lucidrains/soundstorm-pytorch
A PyTorch implementation of SoundStorm, Google DeepMind's efficient parallel audio generation model that applies MaskGiT-style masked gener…
391546active
RunanywhereAI/RCLI
RCLI is a single-binary CLI that runs open-source AI models locally on your machine, covering chat, vision, speech-to-text, text-to-speech,…
821542active
vibevoice-community/VibeVoice
VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker …
601542active
pipecat-ai/smart-turn
Smart Turn is an open-source, audio-native turn detection model that decides when a voice agent should respond to human speech, using proso…
491538active
bytedance/SALMONN
SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video…
731513active
kyutai-labs/hibiki
Hibiki is a decoder-only model for streaming (simultaneous) speech-to-speech translation, built on the multistream Moshi architecture. It p…
281509active
amicalhq/amical
Amical is an open-source, local-first AI dictation app for macOS and Windows that transcribes speech with Whisper and enhances it with LLMs…
841506active
kyutai-labs/unmute
Unmute is a system that lets any text LLM listen and speak by wrapping it with Kyutai's low-latency speech-to-text and text-to-speech model…
601506active
Flashlight
wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into …
625466maintenance
stepfun-ai/Step-Audio2
Step-Audio 2 is an end-to-end multimodal large language model for industry-strength audio understanding and speech conversation, with open-…
501503active
QwenAudio/Fun-ASR
Fun-ASR is a family of open-source LLM-based end-to-end speech recognition models from Tongyi Lab, covering Chinese, dialects, accents, and…
811496active
ibab/tensorflow-wavenet
A TensorFlow implementation of DeepMind's WaveNet generative neural network architecture for raw audio waveform generation. It provides tra…
325428maintenance
espressif/esp-sr
ESP-SR is Espressif's speech recognition framework for ESP32-series chips, providing wake word detection (WakeNet), voice activity detectio…
781492active
ekwek1/soprano
Soprano is an ultra-lightweight 80M-parameter text-to-speech model and Python library for fast, expressive, high-fidelity speech synthesis …
441486active
fossasia/voxbento
Voxbento is an open-source, self-hosted platform for real-time simultaneous interpretation at live events. Interpreters broadcast translate…
771482active
k2-fsa/icefall
Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi…
641482active
IBM Watson Node.js SDK
The official Node.js/TypeScript client library (npm package ibm-watson) for accessing IBM Watson services such as Assistant, Speech to Text…
841473active
McCloudS/subgen
Subgen is a self-hosted service that automatically generates subtitles for media files using OpenAI's Whisper speech recognition models. It…
741465active
NVIDIA/tacotron2
NVIDIA's PyTorch implementation of the Tacotron 2 text-to-speech model, which synthesizes mel spectrograms from text for vocoder-based audi…
325296maintenance
microsoft/NeuralSpeech
NeuralSpeech is a Microsoft Research Asia research repository containing implementations of neural speech processing models across ASR erro…
321462active
silverstein/minutes
Minutes is an open-source, privacy-first conversation memory app that records meetings, voice memos, and dictation, transcribes them locall…
771453active
DicioTeam/dicio-android
Dicio is a free and open source voice assistant app for Android that interprets user questions and provides speech and graphical feedback e…
831452active
jxlpzqc/TMSpeech
TMSpeech is a Windows desktop application that captures system audio via WASAPI loopback and performs real-time Chinese speech-to-text, dis…
821447active
voice-cloning-app/Voice-Cloning-App
A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom…
231440active
edwko/OuteTTS
OuteTTS is a Python interface for running OuteAI's text-to-speech models, supporting multiple inference backends including llama.cpp, Huggi…
521436active
FireRedTeam/FireRedTTS2
FireRedTTS-2 is a long-form streaming text-to-speech system for multi-speaker dialogue generation, built in PyTorch with a dual-transformer…
401428active
devnen/Chatterbox-TTS-Server
A self-hosted server wrapping Resemble AI's Chatterbox TTS models behind an OpenAI-compatible API with a modern web UI. It supports voice c…
651420active
lenML/Speech-AI-Forge
Speech-AI-Forge is a Python project built around multiple TTS generation models (ChatTTS, CosyVoice, Fish-Speech, Index-TTS, F5-TTS, FireRe…
621416active
joewongjc/type4me
Type4Me is a macOS voice input method that transcribes speech in real time (streaming, local or cloud ASR engines) and refines text with LL…
811404active
0nutation/SpeechGPT
SpeechGPT is a series of speech large language models that can perceive and generate speech, supporting cross-modal instruction following a…
291401active
zackees/transcribe-anything
A Python CLI app that transcribes local audio/video files or URLs (YouTube, Rumble, etc.) using multiple Whisper backends with automatic de…
931393active
wenet-e2e/wespeaker
WeSpeaker is a research and production-oriented toolkit for speaker embedding learning, supporting speaker verification, recognition, and d…
641392active
haoheliu/voicefixer
VoiceFixer is a Python library and CLI tool for general speech restoration, using a pretrained neural vocoder to restore degraded human spe…
261373stable
MoonInTheRiver/DiffSinger
Official PyTorch implementation of DiffSinger, an AAAI 2022 paper on singing voice synthesis and text-to-speech using a shallow diffusion m…
654851maintenance
k2-fsa/k2
k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)…
641352active
nyrahealth/CrisperWhisper
CrisperWhisper 2.0 is a controllable speech recognition model and Python library that transcribes audio either verbatim (including fillers,…
891349active
shivammehta25/Matcha-TTS
Matcha-TTS is a PyTorch-based text-to-speech system that uses conditional flow matching for fast, non-autoregressive speech synthesis. It s…
621349active
Softcatala/whisper-ctranslate2
A command-line transcription and translation tool compatible with OpenAI's Whisper CLI, built on CTranslate2 and faster-whisper for up to 4…
601343active
claritylab/lucida
Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros…
324782maintenance
mbailey/voicemode
VoiceMode is a Python-based MCP server and Claude Code plugin that enables natural, hands-free voice conversations with Claude Code and oth…
801339active
mmorise/World
WORLD is a C++ library for high-quality speech analysis, manipulation, and synthesis based on a vocoder design. It estimates F0 (via DIO/Ha…
641338stable
google-gemma/gemma-translator
An on-device, fully offline voice translator application powered by Gemma 4 via LiteRT-LM, with a React web frontend optimized for small ha…
561337active
joey-zhou/xiaozhi-esp32-server-java
A Java enterprise-grade server and management platform for the Xiaozhi ESP32 AI voice assistant hardware, providing a full front-end/back-e…
691336active
GauravSingh9356/J.A.R.V.I.S
A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,…
481332active
andimarafioti/faster-qwen3-tts
A Python library for real-time text-to-speech inference with Qwen3-TTS using manual CUDA graph capture, requiring no Flash Attention, vLLM,…
781330active
sanchit-gandhi/whisper-jax
An optimized JAX implementation of OpenAI's Whisper speech recognition model, built on Hugging Face Transformers, offering up to 70x faster…
304682maintenance
jishengpeng/WavTokenizer
WavTokenizer is a state-of-the-art discrete neural audio codec that compresses speech, music, and general audio into only 40 or 75 discrete…
271316active
yeyupiaoling/VoiceprintRecognition-Pytorch
A PyTorch-based voiceprint recognition (speaker recognition) framework implementing models such as ECAPA-TDNN, ResNetSE, ERes2Net, and CAM+…
581312active
Henry-23/VideoChat
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late…
481303active
stenolabs/stenoai
Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire…
841299active
ictnlp/StreamSpeech
StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p…
371287active

← prev page 3 / 6 next →