Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: speech-processing

552 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
YannickJadoul/Parselmouth
Parselmouth is a Python library that exposes Praat's internal C/C++ speech and acoustic analysis code through a complete, Pythonic interfac…
741286stable
jasperproject/jasper-client
Client code for the Jasper voice computing platform, an open source platform for building always-on, voice-controlled applications. It prov…
324518maintenance
studio-dots-ai/dots.tts
dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive text-to-speech system, distributed as a Python library with pretrain…
791275active
phuc-nt/my-translator
A Tauri-based desktop app that captures system or microphone audio, transcribes it, and shows translations in a minimal real-time overlay, …
761275active
ABexit/ASR-LLM-TTS
An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) …
601271active
yzfly/douyin-mcp-server
A Python MCP server and WebUI that extracts watermark-free Douyin (TikTok China) video download links and transcribes video speech into tex…
101258active
peteonrails/voxtype
Voxtype is a local, offline voice-to-text dictation tool for Linux with push-to-talk hotkey support across Wayland compositors like Hyprlan…
781252active
gitmylo/audio-webui
An all-in-one web UI for audio-related neural networks, bundling text-to-speech (Bark), voice conversion/cloning (RVC), and text-to-audio/m…
301247active
sh-lee-prml/HierSpeechpp
Official PyTorch implementation of HierSpeech++, a fast zero-shot speech synthesizer for text-to-speech and voice conversion based on hiera…
281238active
Ksuriuri/index-tts-vllm
A reimplementation of IndexTTS's GPT model inference using vLLM, providing significantly faster text-to-speech generation with a web UI and…
581235active
NVIDIA/BigVGAN
BigVGAN is NVIDIA's official PyTorch implementation of a universal neural vocoder (ICLR 2023) that generates high-fidelity raw audio wavefo…
231227stable
yeyupiaoling/Whisper-Finetune
A toolkit for fine-tuning OpenAI's Whisper speech recognition models using LoRA, supporting training with or without timestamps and even wi…
661223active
Aratako/Irodori-TTS
Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer…
581219active
modal-labs/quillman
QuiLLMan is a voice chat application built on Kyutai's Moshi speech-to-speech language model, deployed serverlessly on Modal with a FastAPI…
681213active
hgneng/ekho
Ekho is an open-source Chinese text-to-speech engine supporting Mandarin, Cantonese, and Tibetan, part of the eGuideDog accessibility proje…
681211active
aTrainTranscription/aTrain
aTrain is a desktop GUI application for offline transcription of speech recordings using Whisper-based machine learning models, with speake…
791205active
metavoiceio/metavoice-src
MetaVoice-1B is a 1.2B parameter foundational text-to-speech model trained on 100K hours of speech, focused on emotional rhythm and tone in…
264205maintenance
hkjarral/AVA-AI-Voice-Agent-for-Asterisk
An open-source AI voice agent that integrates with Asterisk/FreePBX phone systems via Audiosocket/RTP, built in Python with a modular pipel…
851202active
nishuzumi/gemini-teacher
A Python CLI application that acts as an English speaking practice assistant powered by Google Gemini. It listens to your speech via microp…
681200active
kizuna-ai-lab/sokuji
Sokuji is a cross-platform real-time two-way speech translation app for bilingual meetings, available as a desktop application (Windows, ma…
831199active
diodiogod/TTS-Audio-Suite
A ComfyUI custom node suite providing unified multi-engine Text-to-Speech, Voice Conversion, and audio editing across 19 engines like Chatt…
841185active
digimata/parrot
Parrot is an ultra-minimalist push-to-talk dictation daemon for macOS that transcribes speech entirely on-device using WhisperKit on the Ap…
731185active
goodroot/hyprwhspr
hyprwhspr is a native Linux system-wide speech-to-text dictation application supporting local models (Whisper, Parakeet, Cohere) with optio…
851182active
NVIDIA/audio-flamingo
NVIDIA's PyTorch implementation of the Audio Flamingo series of large audio-language models (AF1, AF2, AF3, and Music Flamingo) for audio u…
501182active
ddlBoJack/emotion2vec
Official PyTorch implementation of emotion2vec, a self-supervised pre-trained model for speech emotion representation. It provides code for…
271179active
nari-labs/dia2
Dia2 is a streaming dialogue text-to-speech model from Nari Labs that can generate conversational audio in realtime without needing the ful…
411172active
baidu-research/warp-ctc
A fast parallel implementation of the Connectionist Temporal Classification (CTC) loss function for CPU and CUDA GPU, with a simple C inter…
324069maintenance
iver56/torch-audiomentations
A PyTorch library for fast audio data augmentation, inspired by audiomentations. It provides GPU-accelerated, differentiable audio transfor…
581167active
soniqo/speech-swift
An open-source Swift toolkit for on-device speech AI on Apple Silicon, providing ASR, TTS, speech-to-speech, VAD, and speaker diarization v…
831161active
gemelo-ai/vocos
Vocos is a fast neural vocoder that synthesizes audio waveforms from acoustic features such as mel-spectrograms or EnCodec tokens. It uses …
651155stable
janvarev/Irene-Voice-Assistant
Irene is an offline-capable Russian voice assistant written in Python that recognizes speech via Vosk and responds using TTS engines. Its f…
651152active
TensorSpeech/TensorFlowTTS
TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me…
233995maintenance
lhotse-speech/lhotse
Lhotse is a Python library for flexible, scalable preparation of multimodal (speech, audio, video, image, text) data for machine learning, …
891149active
google/lyra
Lyra is a very low-bitrate speech codec from Google that combines traditional codec techniques with generative machine learning models to c…
233973maintenance
midas-research/audino
Audino is an open-source, self-hosted web application for annotating audio, supporting transcription, labeling, and speaker-related tasks. …
521146active
facebookresearch/fairseq2
fairseq2 is a PyTorch-based sequence modeling toolkit from Meta FAIR for training custom models for content generation tasks such as langua…
891143active
andabi/deep-voice-conversion
A TensorFlow implementation of deep neural networks for voice conversion (voice style transfer) that converts a source speaker's voice into…
323938maintenance
xzf-thu/Mega-ASR
Mega-ASR is a foundation automatic speech recognition model trained on 2.6M samples spanning 7 atomic acoustic conditions and 54 compound r…
591136active
video-db/call.md
Call.md is an open-source Electron desktop app that records meetings locally, transcribes them in real time with speaker separation, and pr…
591132active
sooftware/conformer
An unofficial PyTorch implementation of the Conformer architecture (convolution-augmented Transformer) from the INTERSPEECH 2020 paper, tar…
721131active
PriesiaMioShirakana/DragonianVoice
A C++ inference library for running ONNX-based TTS, SVC (singing voice conversion), and SVS (singing voice synthesis) models, supporting ar…
401129active
PromtEngineer/Verbi
Verbi is a modular Python voice assistant application for experimenting with state-of-the-art transcription, LLM response generation, and t…
481124active
ardha27/AI-Waifu-Vtuber
An AI waifu Vtuber application that listens to your voice via Whisper speech recognition, generates in-character replies with OpenAI, and s…
681113active
KoljaB/RealtimeVoiceChat
A Python client-server application enabling natural spoken conversations with LLMs, streaming browser audio via WebSockets through Realtime…
333832maintenance
gabber-dev/gabber
Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr…
441111active
vocodedev/vocode-core
Vocode is an open-source Python library for building real-time, voice-based LLM applications and agents. It orchestrates streaming transcri…
213785maintenance
savbell/whisper-writer
WhisperWriter is a small desktop dictation app that transcribes microphone audio to text using OpenAI's Whisper model, either locally via f…
201097active
Saik0s/Whisperboard
WhisperBoard is an open-source iOS app for voice recording and transcription powered by OpenAI's Whisper model running on-device via whispe…
571093active
festvox/flite
Flite (festival-lite) is a small, fast, portable run-time text-to-speech synthesis engine written entirely in ANSI C by Carnegie Mellon Uni…
231081stable
XiaomiMiMo/MiMo-Audio
Xiaomi's open-source 7B audio language model family (Base and Instruct) plus a 1.2B RVQ audio tokenizer, pretrained on 100M+ hours of audio…
571077active
Muesli-HQ/muesli
Muesli is an open-source native macOS app that combines hotkey-driven AI dictation with local meeting transcription, running speech-to-text…
821072active
ufal/whisper_streaming
A Python library that turns Whisper-like speech recognition models into a real-time streaming transcription and translation system using a …
533672maintenance
brenpoly/be-more-agent
An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp…
641060active
Makememo/MemoAI
MemoAI is a desktop application for macOS and Windows that transcribes audio and video (YouTube links, podcasts, local files) into text and…
961059active
X-LANCE/SLAM-LLM
SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc…
551056active
zai-org/GLM-TTS
GLM-TTS is a Python-based text-to-speech synthesis system built on large language models, using a two-stage LLM plus Flow model architectur…
501055active
alphacep/vosk-android-demo
A demo Android application showing offline speech recognition and speaker identification using the Vosk and Kaldi libraries. It serves as a…
481055active
wxxxcxx/ms-ra-forwarder
A self-hostable free online text-to-speech API that proxies Microsoft Edge 'Read Aloud' and Azure TTS demo endpoints. It can be deployed vi…
611053active
Artrajz/vits-simple-api
A Python HTTP API service that exposes VITS-family text-to-speech models (VITS, Bert-VITS2, GPT-SoVITS, W2V2 emotional VITS) for inference.…
691052active
introlab/odas
ODAS (Open embeddeD Audition System) is a C library for real-time sound source localization, tracking, separation, and post-filtering using…
321048stable
Kieirra/murmure
Murmure is a privacy-first, open-source desktop speech-to-text application that transcribes voice entirely on-device using NVIDIA's Parakee…
851047active
shang-zhu/violin
Violin is an open-source video translation tool that transcribes speech, translates it into 33 languages, synthesizes a native-sounding voi…
521047active
k2-fsa/ZipVoice
ZipVoice is a series of fast, high-quality zero-shot text-to-speech models based on flow matching, with a compact 123M-parameter Zipformer-…
431045active
xiaochong/hi-kid
HiKid is a free, open-source Electron desktop app that lets children in non-English-speaking countries practice English speaking and listen…
541043active
PatterAI/Patter
Patter is an open-source, MIT-licensed SDK (Python and TypeScript) that connects AI agents to real phone calls, handling telephony, speech-…
781039active
Goekdeniz-Guelmez/Local-NotebookLM
A local, open-source alternative to Google's NotebookLM that converts PDF documents into audio content like podcasts, summaries, and interv…
741027active
rayenfeng/riko_project
Project Riko is an anime-themed conversational voice assistant that combines OpenAI's GPT for dialogue, GPT-SoVITS for voice synthesis, and…
311023active
hezarai/hezar
Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta…
781013active
TensorSpeech/TensorFlowASR
TensorFlowASR is a Python library implementing automatic speech recognition architectures such as DeepSpeech2, Jasper, RNN Transducer, Cont…
661010active
HumeAI/tada
TADA is an open-source speech-language model from Hume AI that generates expressive, high-fidelity speech via text-acoustic dual alignment,…
511009active
voquill/voquill
Voquill is an open-source, cross-platform AI voice dictation app that lets users dictate into any desktop application, with AI-powered tran…
771008active
AutoArk/open-audio-opd
An industrial training stack for online policy distillation (OPD) of audio models, distilling compact ASR (and planned TTS) student models …
521007active
fossasia/voxbento_translator
A Flask-based HTTP API for real-time speech-to-text transcription and translation that accepts streamed audio chunks and routes them to plu…
751002active
resemble-ai/Resemblyzer
Resemblyzer is a Python package that uses a deep learning voice encoder to convert speech audio into 256-dimensional voice embeddings. Thes…
233300maintenance
linyiLYi/bilibot
A local chatbot fine-tuned from Bilibili user comments, built on Qwen1.5-32B-Chat using Apple's MLX LoRA fine-tuning. It supports text chat…
243139maintenance
pluja/whishper
Whishper is a self-hosted, 100% local audio transcription and subtitling suite with a web UI, powered by FasterWhisper. It transcribes audi…
613066maintenance
keithito/tacotron
An unofficial open-source TensorFlow implementation of Google's Tacotron end-to-end neural text-to-speech model, with a pre-trained model a…
232995maintenance
rishikanthc/Scriberr
Scriberr is an open-source, fully offline audio transcription application designed for self-hosters who value privacy. It transcribes audio…
702990maintenance
pytorch/audio
TorchAudio is PyTorch's audio library providing data manipulation, transforms, and dataset loaders for audio and speech machine learning. I…
872927maintenance
tensorflow/lingvo
Lingvo is a TensorFlow-based framework for building neural networks, particularly sequence models, with a focus on speech recognition, mach…
722864maintenance
readbeyond/aeneas
aeneas is a Python/C library and set of CLI tools that automatically computes forced alignments, generating a synchronization map between a…
652863maintenance
marytts/marytts
MaryTTS is an open-source, multilingual text-to-speech synthesis platform written in pure Java, operating as a client-server system. It sup…
242583maintenance
zzmp/juliusjs
JuliusJS is a JavaScript port of the Julius speech recognition engine that runs entirely in the browser via a Web Worker. It transcribes us…
232565maintenance
s3prl/s3prl
S3PRL is a PyTorch toolkit for self-supervised speech pre-training and representation learning, bundling many upstream models like wav2vec …
552561maintenance
wiseman/py-webrtcvad
A Python wrapper around Google's WebRTC Voice Activity Detector, classifying short frames of 16-bit mono PCM audio as speech or non-speech.…
322496maintenance
jameslyons/python_speech_features
A Python library for extracting common speech features used in automatic speech recognition, including MFCCs, filterbank energies, log filt…
232423maintenance
CjangCjengh/MoeGoe
MoeGoe is an executable command-line tool for running inference with VITS text-to-speech models, supporting TTS, voice conversion, HuBERT-V…
232423maintenance
mravanelli/pytorch-kaldi
PyTorch-Kaldi is a toolkit for developing state-of-the-art DNN/HMM hybrid speech recognition systems, combining PyTorch-managed neural netw…
322404maintenance
r9y9/wavenet_vocoder
A PyTorch implementation of the WaveNet vocoder that generates high-quality raw speech waveforms conditioned on acoustic features like mel-…
232376maintenance
NVIDIA/waveglow
WaveGlow is a PyTorch implementation of a flow-based generative network that synthesizes high-quality speech audio from mel-spectrograms, c…
322339maintenance
Rayhane-mamah/Tacotron-2
A TensorFlow implementation of DeepMind's Tacotron-2 neural text-to-speech architecture, including both the Tacotron spectrogram predictor …
322322maintenance
jianfch/stable-ts
A Python library that modifies OpenAI's Whisper to produce more reliable timestamps, adding transcription, forced alignment, and audio inde…
102281maintenance
fatchord/WaveRNN
A PyTorch implementation of DeepMind's WaveRNN neural vocoder plus a Tacotron text-to-speech system, trained on LJSpeech. It supports train…
322190maintenance
ming024/FastSpeech2
A PyTorch implementation of Microsoft's FastSpeech 2 text-to-speech model, supporting English and Mandarin with single- and multi-speaker s…
322185maintenance
SeanNaren/deepspeech.pytorch
A PyTorch implementation of the DeepSpeech2 speech recognition model, built on PyTorch Lightning, supporting training, testing, and inferen…
232136maintenance
r9y9/deepvoice3_pytorch
A PyTorch implementation of Deep Voice 3 and related convolutional neural network-based text-to-speech synthesis models. It includes traini…
231975maintenance
nobody132/masr
MASR is an end-to-end Mandarin Chinese automatic speech recognition project built on a gated convolutional neural network (similar to Wav2L…
231968maintenance
QwenLM/Qwen-Audio
Official repository for Qwen-Audio, Alibaba Cloud's large audio-language model with pretrained and chat variants. It provides model weights…
271945maintenance
julius-speech/julius
Julius is an open-source large vocabulary continuous speech recognition (LVCSR) decoder written in C, based on word N-gram language models …
351933maintenance
astorfi/lip-reading-deeplearning
A TensorFlow implementation of coupled 3D convolutional neural networks for cross audio-visual matching recognition, accompanying an IEEE A…
231904maintenance

← prev page 4 / 6 next →