Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: speech-processing

552 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
openai/whisper
OpenAI's Whisper is a general-purpose speech recognition model and Python library built on a Transformer sequence-to-sequence architecture.…
68107981stable
RVC-Boss/GPT-SoVITS
GPT-SoVITS is a Python-based few-shot voice cloning and text-to-speech system with an integrated WebUI. It supports zero-shot TTS from a 5-…
6861255active
microsoft/VibeVoice
VibeVoice is Microsoft's open-source frontier voice AI framework combining a next-token diffusion text-to-speech model for expressive, long…
6053225active
ggml-org/whisper.cpp
A high-performance C/C++ port of OpenAI's Whisper automatic speech recognition model with no external dependencies. It supports CPU, GPU (M…
9953203stable
jamiepine/voicebox
Voicebox is a free, open-source, local-first AI voice studio desktop app that combines voice cloning, text-to-speech across 7 engines, and …
7551551active
2noise/ChatTTS
ChatTTS is a generative text-to-speech model optimized for conversational dialogue scenarios such as LLM assistants, supporting English and…
6939797active
myshell-ai/OpenVoice
OpenVoice is a Python library and audio foundation model for instant voice cloning, requiring only a short reference audio clip to replicat…
3437312active
OpenBMB/VoxCPM
VoxCPM is a tokenizer-free text-to-speech system built on a diffusion autoregressive architecture that generates continuous speech represen…
7936145active
fishaudio/fish-speech
Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 l…
7432413active
cjpais/Handy
Handy is a free, open-source, cross-platform desktop application for offline speech-to-text dictation. Press a configurable shortcut, speak…
8330402active
Zackriya-Solutions/meetily
Meetily is a privacy-first, open-source AI meeting assistant that records, transcribes, and summarizes meetings entirely locally using Para…
7929940active
78/xiaozhi-esp32
XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen…
8829186active
resemble-ai/chatterbox
Chatterbox is a family of state-of-the-art open-source text-to-speech models by Resemble AI, including multilingual and low-latency Turbo v…
5226163active
mozilla-ai/llamafile
llamafile is a Mozilla project that packages llama.cpp (and whisper.cpp for whisperfile) with Cosmopolitan Libc into single-file executable…
9425694active
SYSTRAN/faster-whisper
A fast reimplementation of OpenAI's Whisper speech-to-text model built on the CTranslate2 inference engine, offering up to 4x speedup and l…
5725104active
m-bain/whisperX
WhisperX is a Python library providing fast automatic speech recognition with word-level timestamps and speaker diarization. It combines Wh…
9123765active
index-tts/index-tts
IndexTTS is an industrial-level zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest IndexTTS…
7823511active
QwenAudio/CosyVoice
CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-…
6122925active
chidiwilliams/buzz
Buzz is a cross-platform desktop application that transcribes and translates audio and video offline using OpenAI's Whisper models. It offe…
9521147active
DrewThomasson/ebook2audiobook
A Python application that converts non-DRM e-books (e.g., EPUB, PDF) into audiobooks with chapters and metadata using TTS engines like XTTS…
9320053active
modelscope/FunASR
FunASR is an industrial end-to-end speech recognition toolkit built on PyTorch, offering ASR, VAD, punctuation restoration, speaker diariza…
9920036active
nari-labs/dia
Dia is a 1.6B-parameter open-weight text-to-speech model from Nari Labs that generates ultra-realistic multi-speaker dialogue in a single p…
4319378active
NVIDIA-NeMo/Speech
NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, …
9818337active
KittenML/KittenTTS
KittenTTS is an open-source, ultra-lightweight text-to-speech library built on ONNX, with models ranging from 15M to 80M parameters (25-80 …
6615403active
SWivid/F5-TTS
F5-TTS is the official implementation of a fully non-autoregressive text-to-speech system based on flow matching with a Diffusion Transform…
8515167active
Vosk
Vosk is an offline open source speech recognition toolkit supporting 20+ languages with small models and a streaming API. It provides bindi…
6615077stable
pipecat-ai/pipecat
Pipecat is an open-source Python framework (BSD-2) for building real-time voice and multimodal conversational AI agents. It orchestrates sp…
8914769active
SesameAILabs/csm
CSM (Conversational Speech Model) is Sesame's speech generation model that produces conversational audio from text and audio context, using…
3014720active
k2-fsa/sherpa-onnx
sherpa-onnx is an offline speech processing toolkit built on next-gen Kaldi and onnxruntime, supporting speech-to-text, text-to-speech, spe…
9514411active
livekit/agents
LiveKit Agents is an open-source framework for building realtime, multimodal voice AI agents that run as programmable participants in LiveK…
9013176active
QwenLM/Qwen3-TTS
Qwen3-TTS is a series of open-source text-to-speech models from Alibaba's Qwen team, supporting expressive and streaming speech generation,…
4813113active
Vaibhavs10/insanely-fast-whisper
A CLI tool for transcribing audio files on-device using OpenAI's Whisper models, powered by Hugging Face Transformers, Optimum, and Flash A…
4913049active
huggingface/speech-to-speech
A low-latency, fully modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) built from open-source models, exposed via the OpenAI Realtime…
8512882active
PaddlePaddle/PaddleSpeech
PaddleSpeech is an open-source speech and audio toolkit built on the PaddlePaddle deep learning platform, covering ASR with punctuation, st…
6612670active
abus-aikorea/voice-pro
Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, …
8412647active
facebookresearch/seamless_communication
A library of foundational multilingual multimodal AI models from Meta for speech and text translation, including SeamlessM4T, SeamlessExpre…
7111842active
rany2/edge-tts
A Python module and CLI that lets you use Microsoft Edge's online neural text-to-speech service without needing Edge, Windows, or an API ke…
7911803active
speechbrain/speechbrain
SpeechBrain is an open-source PyTorch-based speech toolkit for building conversational AI systems. It provides training recipes, pretrained…
8311785active
debpalash/VoiceStudio
VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, vid…
8011737active
CorentinJ/Real-Time-Voice-Cloning
A Python implementation of the SV2TTS (Transfer Learning from Speaker Verification to Multispeaker TTS) framework that clones a voice from …
6460110maintenance
TEN-framework/ten-framework
TEN is an open-source framework for building real-time multimodal conversational AI agents, with a focus on low-latency voice assistants. I…
8511082active
SparkAudio/Spark-TTS
Spark-TTS is an LLM-based text-to-speech system built on Qwen2.5 that generates speech directly from single-stream decoupled speech tokens.…
2711006active
QuentinFuxa/WhisperLiveKit
WhisperLiveKit is a self-hosted, ultra-low-latency real-time speech-to-text pipeline built on state-of-the-art simultaneous speech research…
8710962active
kyutai-labs/moshi
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai, built around the Mimi streaming neural audio…
6210949active
altic-dev/FluidVoice
FluidVoice is an open-source (GPLv3) macOS dictation app that transcribes speech to text entirely on-device using models like Parakeet, Whi…
8410946active
moonshine-ai/moonshine
Moonshine Voice is an open-source on-device AI toolkit providing very low latency speech-to-text, intent recognition, and text-to-speech fo…
8910944active
pyannote/pyannote-audio
pyannote.audio is an open-source Python toolkit built on PyTorch for speaker diarization, providing neural building blocks like voice activ…
9510475active
xinnan-tech/xiaozhi-esp32-server
A self-hosted backend server for the xiaozhi-esp32 open-source smart hardware project, implementing the Xiaozhi communication protocol over…
8110433active
NVIDIA/personaplex
PersonaPlex is a real-time, full-duplex speech-to-speech conversational model from NVIDIA that supports persona control via text role promp…
4710393active
nexmoe/VidBee
VidBee is a free, open-source Electron desktop app for downloading video and audio from 1000+ sites (via yt-dlp) and importing local media.…
8410386active
open-mmlab/Amphion
Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s…
5110271active
KoljaB/RealtimeSTT
RealtimeSTT is a Python library for low-latency, real-time speech-to-text with voice activity detection, wake word activation, and streamin…
9410080active
snakers4/silero-vad
Silero VAD is a pre-trained, enterprise-grade Voice Activity Detector model available via PyPI, runnable with PyTorch or ONNX Runtime. It d…
8610061stable
espnet/espnet
ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance…
889941active
k2-fsa/OmniVoice
OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style a…
799455active
coqui-ai/TTS
Coqui TTS is a deep learning toolkit for text-to-speech synthesis, providing pretrained models in over 1100 languages plus tools for traini…
2345953maintenance
QwenAudio/SenseVoice
SenseVoice is an open-source speech foundation model (SenseVoiceSmall) providing multilingual ASR, spoken language identification, speech e…
889151active
kyutai-labs/pocket-tts
Pocket TTS is a lightweight 100M-parameter text-to-speech model and Python library from Kyutai that runs efficiently on CPUs without a GPU.…
839135active
jianchang512/clone-voice
A voice cloning tool with a web interface built on the coqui.ai xtts_v2 model, letting users synthesize speech in any voice from text or co…
108989active
Uberi/speech_recognition
A Python library for performing speech recognition with support for multiple engines and APIs, both online (Google, Azure, Wit.ai, OpenAI W…
948987active
jasonppy/VoiceCraft
VoiceCraft is a token infilling neural codec language model for zero-shot speech editing and text-to-speech on in-the-wild data like audiob…
638574active
hexgrad/kokoro
An inference library for Kokoro-82M, an open-weight text-to-speech model with 82 million parameters that delivers quality comparable to lar…
368570active
netease-youdao/EmotiVoice
EmotiVoice is an open-source text-to-speech engine supporting English and Chinese with over 2000 voices and prompt-controlled emotional syn…
178523active
santinic/audiblez
Audiblez is a Python CLI tool (with an optional GUI) that converts .epub e-books into .m4b audiobooks using the Kokoro-82M text-to-speech m…
528460active
nl8590687/ASRT_SpeechRecognition
ASRT is a deep-learning-based Chinese speech recognition (speech-to-text) system built with TensorFlow/Keras, using CNN, LSTM, attention me…
578383active
GetStream/Vision-Agents
An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O…
848100active
suno-ai/bark
Bark is Suno's open-source transformer-based text-to-audio model that generates highly realistic multilingual speech, music, background noi…
3039249maintenance
Blaizzy/mlx-audio
MLX-Audio is a Python library built on Apple's MLX framework for fast text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS…
887793active
jianchang512/ChatTTS-ui
A local web interface for the ChatTTS text-to-speech model that synthesizes speech from mixed Chinese/English text, numbers, and symbols. I…
707640active
babysor/MockingBird
MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real…
5436909maintenance
myshell-ai/MeloTTS
MeloTTS is a high-quality multi-lingual text-to-speech library supporting English (multiple accents), Spanish, French, Chinese, Japanese, a…
167608active
Zyphra/Zonos
Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I…
247244active
thewh1teagle/vibe
Vibe is a cross-platform desktop app for fully offline audio and video transcription using Whisper, Nemotron, and Parakeet models. It suppo…
937206active
wzpan/wukong-robot
wukong-robot is a modular Chinese-language voice assistant / smart speaker project in Python that combines offline wake-word detection, ASR…
327126active
openai/openai-realtime-agents
A Next.js TypeScript demo application showcasing advanced agentic patterns (Chat-Supervisor and Sequential Handoff) built on the OpenAI Rea…
486966active
PaddlePaddle/models
PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda…
236932active
TalAter/annyang
annyang is a tiny (2 KB), dependency-free JavaScript library that adds speech recognition and voice commands to websites using the Web Spee…
786816stable
espeak-ng/espeak-ng
eSpeak NG is a compact open-source text-to-speech synthesizer supporting over 100 languages and accents, using formant synthesis for small …
676763active
HaujetZhao/CapsWriter-Offline
CapsWriter-Offline is a fully offline voice input tool for Windows that transcribes speech to text when you hold CapsLock or mouse side but…
926691active
argmaxinc/argmax-oss-swift
A Swift SDK providing turn-key on-device speech AI frameworks for Apple Silicon, including WhisperKit (speech-to-text with Whisper), Speake…
916338active
yl4579/StyleTTS2
StyleTTS 2 is a PyTorch text-to-speech model that uses style diffusion and adversarial training with large speech language models (e.g., Wa…
296336active
canopyai/Orpheus-TTS
Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a…
456314active
neuphonic/neutts
NeuTTS is a collection of open-source, on-device text-to-speech models built on small LLM backbones, with instant voice cloning from as lit…
606256active
tyiannak/pyAudioAnalysis
pyAudioAnalysis is a Python library for audio analysis covering feature extraction (MFCCs, spectrograms, chromagrams), supervised and unsup…
486254active
modelscope/FunClip
FunClip is an open-source, locally deployed video clipping tool that uses FunASR Paraformer models for speech recognition and subtitle gene…
956190active
Beingpax/VoiceInk
VoiceInk is a native macOS voice-to-text dictation app that transcribes speech to text almost instantly using local AI models (Parakeet, Wh…
846115active
bytedance/MegaTTS3
MegaTTS 3 is ByteDance's open-source PyTorch text-to-speech model with a lightweight 0.45B-parameter Diffusion Transformer backbone. It pro…
596091active
xiph/rnnoise
RNNoise is a C library that uses a hybrid DSP/recurrent neural network approach for real-time full-band speech noise suppression. It also s…
265801stable
OpenWhispr/openwhispr
OpenWhispr is an open-source, privacy-first voice-to-text dictation desktop app for macOS, Windows, and Linux. It supports fully local tran…
815766active
dnhkng/GLaDOS
A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,…
645689active
MahmoudAshraf97/whisper-diarization
A pipeline that combines OpenAI Whisper transcription with speaker diarization using Voice Activity Detection (MarbleNet) and speaker embed…
755630active
xiangyuecn/Recorder
A JavaScript HTML5 audio recording library for browsers and hybrid apps, supporting mp3, wav, pcm, ogg, amr, webm, and g711 formats with re…
845626active
huggingface/parler-tts
Parler-TTS is a lightweight text-to-speech library from Hugging Face that generates high-quality, natural-sounding speech controllable via …
255586active
dograh-hq/dograh
Dograh is an open-source, self-hostable voice AI platform for building production voice agents, positioned as an alternative to Vapi and Re…
845505active
modstart-lib/aigcpanel
AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, …
885482active
remsky/Kokoro-FastAPI
A Dockerized FastAPI wrapper around the Kokoro-82M text-to-speech model exposing an OpenAI-compatible speech endpoint with CPU, NVIDIA, AMD…
885373active
ysharma3501/LuxTTS
LuxTTS is a lightweight zipvoice-based text-to-speech model for high-quality zero-shot voice cloning, generating clear 48kHz speech at up t…
545306active
OHF-Voice/piper1-gpl
Piper is a fast, fully local neural text-to-speech engine that embeds espeak-ng for phonemization and ships with a CLI, HTTP web server, Py…
865287active
wenet-e2e/wenet
WeNet is a production-first, end-to-end automatic speech recognition (ASR) toolkit built on PyTorch with transformer/conformer models. It p…
625227active
Plachtaa/VITS-fast-fine-tuning
A Python pipeline for fast fine-tuning of VITS text-to-speech models, enabling speaker adaptation in under an hour from short audio, long a…
105012active

page 1 / 6 next →