Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: speech-processing

552 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
facebookresearch/denoiser
A PyTorch library implementing a causal, real-time speech enhancement model that operates on raw waveforms to remove background noise and r…
101899maintenance
Kyubyong/tacotron
A heavily documented TensorFlow implementation of Tacotron, a fully end-to-end text-to-speech synthesis model. It includes training, prepro…
321832maintenance
kalliope-project/kalliope
Kalliope is a modular, always-on voice-controlled personal assistant framework written in Python, designed to run on Linux systems includin…
231772maintenance
strob/gentle
Gentle is a robust yet lenient forced aligner built on Kaldi that aligns audio speech with a known text transcript. It can be used as a Mac…
651705maintenance
kan-bayashi/ParallelWaveGAN
Unofficial PyTorch implementations of non-autoregressive neural vocoders including Parallel WaveGAN, MelGAN, Multi-band MelGAN, HiFi-GAN, a…
231645maintenance
Delta-ML/delta
DELTA is a deep learning based end-to-end natural language and speech processing platform built on TensorFlow and Python 3. It provides one…
101607maintenance
vanshg/MacAssistant
MacAssistant is a macOS application that integrates the Google Assistant into the Mac menu bar using the Google Assistant SDK. It is writte…
231603maintenance
google/uis-rnn
A Python library implementing the Unbounded Interleaved-State Recurrent Neural Network (UIS-RNN) algorithm for segmenting and clustering se…
101588maintenance
Marak/say.js
A Node.js library that provides text-to-speech by shelling out to platform-native TTS engines (macOS `say`, Windows SAPI, Linux Festival). …
321531maintenance
beyondcode/writeout.ai
A self-hosted Laravel web application that transcribes uploaded audio files using OpenAI's Whisper API and translates the transcripts via t…
711525maintenance
syl22-00/pocketsphinx.js
PocketSphinx.js is a speech recognition library that runs entirely in the web browser, built by compiling the PocketSphinx C recognizer to …
321508maintenance
google/live-transcribe-speech-engine
Android client libraries from Google's Live Transcribe app for streaming real-time speech recognition via the Google Cloud Speech API. It p…
101499maintenance
s-macke/SAM
SAM (Software Automatic Mouth) is a tiny text-to-speech synthesizer written in C, adapted from the 1982 Commodore 64 speech software. It in…
321497maintenance
wit-ai/pywit
pywit is the official Python SDK for Wit.ai, Facebook's natural language processing platform. It provides a Wit client class for extracting…
671485maintenance
YuanGongND/ast
Official PyTorch implementation of the Audio Spectrogram Transformer (AST) from the Interspeech 2021 paper, which applies a Vision Transfor…
321472maintenance
microsoft/SpeechT5
Microsoft's unified-modal speech-text pre-training framework implementing SpeechT5 and related models (SpeechLM, SpeechUT, VALL-E X, WavLLM…
321449maintenance
m1guelpf/yt-whisper
A Python CLI tool that downloads YouTube videos with yt-dlp and generates subtitle files (VTT) using OpenAI's Whisper speech recognition mo…
321446maintenance
cmusphinx/sphinx4
Sphinx-4 is a speaker-independent, continuous speech recognition library written entirely in Java, originating from CMU and industry resear…
321437maintenance
hungtraan/FacebookBot
A Facebook Messenger chatbot ('Optimist Prime') that supports voice recognition, natural language processing, and contextual follow-up conv…
321424maintenance
DragonComputer/Dragonfire
Dragonfire is an open-source virtual assistant for Ubuntu-based Linux distributions, combining speech recognition, text-to-speech, and NLP …
231407maintenance
innnky/emotional-vits
Emotional VITS is a voice synthesis (TTS) model based on VITS that enables emotion-controllable speech generation without requiring manual …
321392maintenance
LuckyHookin/edge-TTS-record
A Windows desktop tool that records Microsoft Edge's online neural text-to-speech voices (e.g., Xiaoxiao, Yunyang) and saves the output as …
231370maintenance
Jackywine/Bella
Bella is a self-hosted Node.js web application that acts as a personalized AI digital companion with voice interaction. It combines Whisper…
486379experimental
kripken/speak.js
speak.js is a port of the eSpeak C++ speech synthesizer to JavaScript via Emscripten, enabling text-to-speech in the browser using only Jav…
321335maintenance
CSTR-Edinburgh/merlin
Merlin is a toolkit from the University of Edinburgh's CSTR for building deep neural network models for statistical parametric speech synth…
231321maintenance
facebookresearch/svoice
SVoice is a PyTorch implementation of the ICML paper 'Voice Separation with an Unknown Number of Multiple Speakers' from Facebook AI Resear…
101315maintenance
Renovamen/Speech-Emotion-Recognition
A Python library implementing speech emotion recognition with Keras/TensorFlow 2 using LSTM, CNN, SVM, and MLP models. It extracts audio fe…
321314maintenance
kakaobrain/pororo
PORORO is a Python library from Kakao Brain that provides unified access to dozens of pretrained neural models for natural language process…
101305maintenance
elanmart/cbp-translate
A demo application that live-translates foreign-language speech in videos into subtitles, mimicking the Cyberpunk 2077 translation effect. …
321274maintenance
sdkcarlos/artyom.js
Artyom.js is a JavaScript library that wraps the Web Speech APIs (webkitSpeechRecognition and speechSynthesis) to add voice control, speech…
231269maintenance
mravanelli/SincNet
SincNet is a PyTorch neural architecture that processes raw audio waveforms using parametrized sinc band-pass filters in the first convolut…
321243maintenance
pollen-robotics/dtw
A small Python module implementing Dynamic Time Warping (DTW), a similarity measure between temporal sequences. It offers a basic pure-Pyth…
231230maintenance
PlayVoice/vits_chinese
A Chinese text-to-speech library combining VITS with BERT-based prosody embeddings and NaturalSpeech infer-loss features, supporting ONNX c…
231227maintenance
Alexander-H-Liu/End-to-end-ASR-Pytorch
A PyTorch implementation of end-to-end automatic speech recognition (ASR), formerly known as Listen, Attend and Spell. It supports seq2seq …
321208maintenance
2fps/recorder
A JavaScript library for recording audio in the browser using the HTML5 Web Audio API. It supports recording, pausing, resuming, playback, …
231206maintenance
clovaai/voxceleb_trainer
A PyTorch framework for training and evaluating speaker recognition and verification models on the VoxCeleb datasets. It implements multipl…
671175maintenance
YaoFANGUK/video-subtitle-generator
A Python application with both GUI and CLI interfaces that generates subtitle files (SRT) from video or audio using local Whisper-based spe…
321175maintenance
bawangxx/XZVoice
XZVoice is a free, open-source desktop text-to-speech application built with Electron, Vue, and ElementUI. It uses Alibaba Cloud's speech s…
231168maintenance
spring-media/TransformerTTS
A TensorFlow 2 implementation of a non-autoregressive Transformer-based neural network for text-to-speech synthesis, based on FastSpeech an…
101161maintenance
Kyubyong/dc_tts
A TensorFlow implementation of DC-TTS, a text-to-speech model based on deep convolutional networks with guided attention. It includes train…
321156maintenance
openinterpreter/01
An open-source voice interface platform that lets users control computers conversationally, powered by Open Interpreter. It pairs a Python …
165156experimental
synesthesiam/opentts
OpenTTS is a text-to-speech server that unifies access to multiple open source TTS systems (Larynx, Glow-Speak, Coqui-TTS, MaryTTS, flite, …
101118maintenance
synesthesiam/voice2json
voice2json is a collection of command-line tools for offline speech-to-text and intent recognition on Linux, supporting 18 languages via en…
101105maintenance
auspicious3000/autovc
AUTOVC is a PyTorch implementation of a many-to-many non-parallel voice conversion framework that performs zero-shot voice style transfer u…
321100maintenance
alumae/kaldi-gstreamer-server
A real-time full-duplex speech recognition server built on the Kaldi toolkit and GStreamer framework, implemented in Python. It streams aud…
321093maintenance
Edresson/YourTTS
YourTTS is a zero-shot multi-speaker text-to-speech and voice conversion model built on VITS, implemented in the Coqui TTS framework. It su…
231053maintenance
aliutkus/speechmetrics
A Python library that wraps several objective speech quality metrics (MOSNet, BSSEval, STOI, PESQ, SRMR, SISDR) behind a unified API. It su…
321051maintenance
pykaldi/pykaldi
PyKaldi is a Python scripting layer providing wrappers for the C++ APIs of the Kaldi speech recognition toolkit and OpenFst library. It ena…
471039maintenance
descriptinc/melgan-neurips
Official PyTorch implementation of MelGAN, a GAN-based non-autoregressive vocoder that inverts mel-spectrograms into raw audio waveforms fo…
321039maintenance
ggeop/Python-ai-assistant
Jarvis is a Python voice-controlled AI assistant for Linux that recognizes speech, responds conversationally, and executes commands like op…
231014maintenance
NATSpeech/NATSpeech
A PyTorch framework for non-autoregressive text-to-speech (NAR-TTS), containing official implementations of PortaSpeech (NeurIPS 2021) and …
231004maintenance
enhuiz/vall-e
An unofficial PyTorch implementation of the VALL-E text-to-speech audio language model, built on the EnCodec tokenizer. It provides trainin…
312976experimental
Priler/jarvis
JARVIS is an offline, privacy-respecting voice assistant built in Rust with Tauri, using neural networks for speech-to-text, text-to-speech…
532909experimental
fikrikarim/parlor
Parlor is a fully on-device, real-time multimodal voice assistant similar to GPT-Live, combining speech recognition, a Gemma vision-languag…
782039experimental
Mini-Omni
Mini-Omni is an open-source multimodal large language model that performs real-time end-to-end speech-to-speech conversation with streaming…
221922experimental
Standard-Intelligence/hertz-dev
Hertz-dev is an open-source 8.5B parameter autoregressive base model for full-duplex conversational audio, released by Standard Intelligenc…
221799experimental
antirez/voxtral.c
A pure C, zero-dependency inference implementation of Mistral's Voxtral Realtime 4B speech-to-text model, with a streaming C API and CLI fo…
451738experimental
collabora/WhisperFusion
WhisperFusion is a real-time voice chat application that combines WhisperLive speech-to-text, a Mistral/Phi LLM, and WhisperSpeech text-to-…
261647experimental
lucidrains/naturalspeech2-pytorch
A PyTorch implementation of NaturalSpeech 2, a zero-shot text-to-speech and singing synthesizer that combines a neural audio codec with a l…
201333experimental
linyiLYi/voice-assistant
A simple single-script Python demo of a local voice assistant that uses Whisper (via Apple MLX) for speech recognition and a local Yi large…
261323experimental
CSCB/vibe-mouse
An open-source Python desktop application that redefines mouse interaction by letting users bind customizable 'Skills' to buttons, with mul…
551086experimental
openai/openai-realtime-api-beta
A Node.js and browser reference client library for OpenAI's Realtime API, enabling real-time voice and text conversations with GPT models. …
221016experimental
susiai/susi_shell
A suite of Python-based command line tools for interacting with AI services directly from the terminal, including chat, text completion, tr…
411002experimental
mozilla/DeepSpeech
DeepSpeech is an open-source, offline speech-to-text engine based on Baidu's Deep Speech research paper and implemented with TensorFlow. It…
1026771abandoned
supertone-inc/supertonic
Supertonic is a lightning-fast, on-device multilingual text-to-speech system powered by ONNX Runtime, with a compact 99M-parameter open-wei…
5813734abandoned
voicepaw/so-vits-svc-fork
A fork of so-vits-svc providing singing voice conversion with realtime support and an improved interface, built on PyTorch and PyTorch Ligh…
829327abandoned
MycroftAI/mycroft-core
Mycroft Core is the core software of the Mycroft open-source voice assistant platform, providing wake-word listening, speech recognition, n…
106611abandoned
unCaptcha
unCaptcha2 is a Python-based security research tool that defeats Google's ReCaptcha v2 audio challenges by submitting the audio to free spe…
324917abandoned
agermanidis/autosub
Autosub is a Python command-line utility that auto-generates subtitles for video or audio files. It performs voice activity detection, tran…
324191abandoned
buriburisuri/speech-to-text-wavenet
A TensorFlow implementation of end-to-end English speech recognition based on DeepMind's WaveNet architecture, trained with CTC loss on sen…
324002abandoned
innnky/so-vits-svc
A singing voice conversion (SVC) framework that uses a SoftVC content encoder with VITS to transform one singer's voice into another timbre…
103779abandoned
Kitt-AI/snowboy
Snowboy is a C++-based hotword (wake word) detection library by KITT.AI with bindings for Python, Android, and other platforms, enabling al…
233364abandoned
zzw922cn/Automatic_Speech_Recognition
An end-to-end automatic speech recognition system implemented in TensorFlow, supporting Mandarin and English with models like DeepSpeech2, …
322831abandoned
coqui-ai/STT
Coqui STT is an open-source deep learning toolkit for training and deploying speech-to-text models, built on TensorFlow with bindings for m…
232606abandoned
react-native-voice/voice
A React Native speech-to-text library providing voice recognition on iOS and Android with both online and offline support. The package is n…
102162abandoned
joshnewlan/say_what
A Python script that listens to conference call audio via speech-to-text (IBM Watson) and alerts the user on HipChat when their name is men…
322087abandoned
wzpan/dingdang-robot
Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi. It uses pluggable STT/TTS en…
101873abandoned
tomlepaine/fast-wavenet
A Python/TensorFlow library implementing an efficient O(L) generation algorithm for Wavenet-style autoregressive models using dynamic progr…
321772abandoned
justLV/onju-voice
A hackable AI home assistant platform that replaces the internals of a Google Nest Mini with a custom ESP32-S3 PCB, paired with a server th…
631711abandoned
Ayanaminn/N46Whisper
A Google Colab notebook application that generates Japanese subtitle files from video using the faster-whisper speech recognition model. It…
361708abandoned
NVIDIA/OpenSeq2Seq
OpenSeq2Seq is a TensorFlow-based toolkit for building and training sequence-to-sequence models for neural machine translation, speech reco…
101558abandoned
elevenlabs/elevenlabs-mcp
The official ElevenLabs Model Context Protocol (MCP) server, exposing ElevenLabs text-to-speech, speech-to-text, voice design, and conversa…
101532abandoned
fossasia/susi_smart_box
SUSI.AI Smart Box is an open-source smart speaker / voice assistant hardware project built around the SUSI.AI assistant. It provides the so…
101529abandoned
sc0ty/subsync
A C++ tool that automatically synchronizes subtitle files with movie or TV audio using speech recognition. It detects spoken audio in the v…
101423abandoned
zenorocha/voice-elements
A pair of Polymer-based Web Components (<voice-player> and <voice-recognition>) that wrap the Web Speech API for speech synthesis (text to …
231348abandoned
dingdang-robot/dingdang-robot
Dingdang is an open-source Chinese voice conversation robot / smart speaker project that runs on Raspberry Pi and other Linux hosts. It use…
101330abandoned
MycroftAI/mimic3
Mimic 3 is a fast, local neural text-to-speech engine developed by Mycroft for the Mark II voice assistant, usable as a Python library, CLI…
291264abandoned
alexa/avs-device-sdk
The Alexa Voice Service (AVS) Device SDK is a C++ SDK for commercial device makers to integrate Alexa voice assistant capabilities directly…
101252abandoned
rhasspy/wyoming-satellite
A Python application that turns a Raspberry Pi (or similar Linux device) with a microphone and speaker into a remote voice satellite using …
101244abandoned
xiph/LPCNet
LPCNet is a low-complexity C implementation of the WaveRNN-based LPCNet neural vocoder for efficient speech synthesis and compression. It a…
321221abandoned
maum-ai/voicefilter
An unofficial PyTorch implementation of Google AI's VoiceFilter system, which isolates a target speaker's voice from noisy mixed audio give…
321216abandoned
basveeling/wavenet
A Keras implementation of DeepMind's WaveNet, a generative neural network model for raw audio synthesis. It supports training on datasets l…
321052abandoned
BrasD99/HeyGenClone
An open-source Python application that clones the HeyGen video translation system, translating videos into multiple languages with voice ov…
101027abandoned
huggingface/transformers
Hugging Face Transformers is a Python library that serves as the model-definition framework for state-of-the-art machine learning models ac…
95164475stable
Unsloth
Unsloth is a desktop application for running and fine-tuning LLMs, diffusion, embedding, and audio models locally, with support for NVIDIA,…
9474883active
mudler/LocalAI
LocalAI is an open-source, self-hosted AI inference runtime that runs LLMs, vision, speech, image, and video models behind OpenAI/Anthropic…
9348696active
huggingface/candle
Candle is a minimalist machine learning framework for Rust focused on performance and ease of use, with CPU and CUDA GPU support. It ships …
7320955active
xszyou/Fay
Fay is an open-source Python digital human framework that connects 2.5D/3D/mobile/web digital humans or OpenAI-compatible LLMs to business …
7513455active
RunanywhereAI/runanywhere-sdks
RunAnywhere is a set of cross-platform SDKs (Swift, Kotlin, React Native, Flutter, TypeScript, C++) over a shared C++ core for running AI m…
8510282active
xorbitsai/inference
Xinference is an open-source model serving platform for deploying LLMs, embedding, speech, image, and multimodal models via a unified OpenA…
919523active

← prev page 5 / 6 next →