Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: speech-recognition

801 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
espeak-ng/espeak-ng
eSpeak NG is a compact open-source text-to-speech synthesizer supporting over 100 languages and accents, using formant synthesis for small …
676763active
HaujetZhao/CapsWriter-Offline
CapsWriter-Offline is a fully offline voice input tool for Windows that transcribes speech to text when you hold CapsLock or mouse side but…
926691active
steipete/summarize
Summarize is a Node.js CLI and Chrome/Firefox extension that extracts clean text from web pages, PDFs, YouTube videos, podcasts, and audio/…
826577active
microsoft/call-center-ai
An AI-powered call center service that lets you initiate or receive phone calls handled by an LLM-driven agent via a simple API call. Built…
646563active
souzatharsis/podcastfy
Podcastfy is an open-source Python package and CLI that transforms multimodal content (websites, PDFs, images, YouTube videos, topics) into…
606521active
argmaxinc/argmax-oss-swift
A Swift SDK providing turn-key on-device speech AI frameworks for Apple Silicon, including WhisperKit (speech-to-text with Whisper), Speake…
916338active
canopyai/Orpheus-TTS
Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a…
456314active
neuphonic/neutts
NeuTTS is a collection of open-source, on-device text-to-speech models built on small LLM backbones, with instant voice cloning from as lit…
606256active
Shaunwei/RealChar
RealChar is an open-source application for creating, customizing, and talking to AI characters/companions in realtime via voice or text. It…
486213active
modelscope/FunClip
FunClip is an open-source, locally deployed video clipping tool that uses FunASR Paraformer models for speech recognition and subtitle gene…
956190active
Beingpax/VoiceInk
VoiceInk is a native macOS voice-to-text dictation app that transcribes speech to text almost instantly using local AI models (Parakeet, Wh…
846115active
bytedance/MegaTTS3
MegaTTS 3 is ByteDance's open-source PyTorch text-to-speech model with a lightweight 0.45B-parameter Diffusion Transformer backbone. It pro…
596091active
cactus-compute/cactus
Cactus is a hybrid edge-cloud AI inference engine for mobile devices, wearables, smart home devices, and robots, built in C++ with custom q…
865934active
joeseesun/qiaomu-anything-to-notebooklm
A Claude Code Skill that ingests content from 15+ sources (WeChat articles, web pages, YouTube, PDFs, EPUB, Office docs, audio) and uploads…
625824active
xiph/rnnoise
RNNoise is a C library that uses a hybrid DSP/recurrent neural network approach for real-time full-band speech noise suppression. It also s…
265801stable
denizsafak/abogen
Abogen is a desktop text-to-speech application that converts EPUB, PDF, text, markdown, and subtitle files into audiobooks with synchronize…
755771active
OpenWhispr/openwhispr
OpenWhispr is an open-source, privacy-first voice-to-text dictation desktop app for macOS, Windows, and Linux. It supports fully local tran…
815766active
dnhkng/GLaDOS
A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,…
645689active
ConnectAI-E/feishu-openai
A self-hosted Go application that integrates OpenAI models (GPT-4, GPT-4V, DALL·E-3, Whisper) into Feishu/Lark as a chatbot. It supports vo…
355638active
MahmoudAshraf97/whisper-diarization
A pipeline that combines OpenAI Whisper transcription with speaker diarization using Voice Activity Detection (MarbleNet) and speaker embed…
755630active
xiangyuecn/Recorder
A JavaScript HTML5 audio recording library for browsers and hybrid apps, supporting mp3, wav, pcm, ogg, amr, webm, and g711 formats with re…
845626active
sohzm/cheating-daddy
Cheating Daddy is a free, open-source Electron desktop app that acts as a real-time AI assistant during video calls, interviews, and meetin…
785580active
father-bot/chatgpt_telegram_bot
A self-hostable Telegram bot that brings ChatGPT and Claude models into Telegram using your own OpenAI/Anthropic/OpenRouter API keys. It su…
785531active
dograh-hq/dograh
Dograh is an open-source, self-hostable voice AI platform for building production voice agents, positioned as an alternative to Vapi and Re…
845505active
liuzhao1225/YouDub-webui
YouDub WebUI is an open-source AI video localization and dubbing tool that converts YouTube, Bilibili, or local videos into target-language…
755358active
OHF-Voice/piper1-gpl
Piper is a fast, fully local neural text-to-speech engine that embeds espeak-ng for phonemization and ships with a CLI, HTTP web server, Py…
865287active
wenet-e2e/wenet
WeNet is a production-first, end-to-end automatic speech recognition (ASR) toolkit built on PyTorch with transformer/conformer models. It p…
625227active
Picovoice/porcupine
Porcupine is a highly-accurate, lightweight on-device wake word detection engine powered by deep neural networks. It enables always-listeni…
734922stable
buxuku/SmartSub
SmartSub (妙幕) is a free, open-source cross-platform desktop application that provides an end-to-end subtitle and dubbing pipeline: speech-t…
894778active
jianchang512/stt
An offline local speech-to-text tool based on faster-whisper models that transcribes audio and video files into JSON, SRT subtitles, or pla…
454766active
Anil-matcha/AI-Youtube-Shorts-Generator
An open-source Python tool that turns long-form YouTube videos into vertical 9:16 short clips using LLM-based highlight detection, Whisper …
664738active
MoonshotAI/Kimi-Audio
Kimi-Audio is an open-source audio foundation model (7B parameters) that unifies audio understanding, generation, and speech conversation i…
314731active
gradio-app/fastrtc
FastRTC is a Python library that turns any Python function into a real-time audio and video stream over WebRTC or WebSockets. It includes b…
584622active
fixie-ai/ultravox
Ultravox is a fast multimodal LLM that understands human speech directly without a separate ASR stage, by projecting audio into the LLM's e…
444555active
ibttf/interview-coder
Interview Coder is an Electron-based desktop application that provides AI assistance during technical interviews, using GPT/OpenAI models t…
454447active
modelscope/ClearerVoice-Studio
ClearerVoice-Studio is an open-source, AI-powered speech processing toolkit from ModelScope/Alibaba offering state-of-the-art pretrained mo…
384445active
OptiKey/OptiKey
OptiKey is a free, open-source on-screen keyboard application for Windows that enables full computer control and speech output using eye-tr…
594411active
cmusphinx/pocketsphinx
PocketSphinx is Carnegie Mellon's lightweight, speaker-independent continuous speech recognition engine, available as a C library, a comman…
844335active
solidSpoon/DashPlayer
DashPlayer is an open-source desktop video player built specifically for English learners, using videos as immersive learning material. It …
994322active
leetcode-mafia/cheetah
Cheetah is a macOS app that provides real-time AI coaching for software engineering interviews. It transcribes conversation audio locally w…
214257active
collabora/WhisperLive
WhisperLive is a nearly-live speech-to-text application built on OpenAI's Whisper, serving real-time transcription over WebSocket and REST …
864241active
openutau/OpenUtau
OpenUtau is a free, open-source singing voice synthesis editor and modern successor to UTAU, built for the UTAU voicebank community. It sup…
704231active
hcfyapp/crx-selection-translate
Huaci Fanyi (Selection Translate) is a browser extension for Chrome, Edge, and Firefox that translates selected text, full web pages, scree…
324139active
huggingface/distil-whisper
Distil-Whisper is a distilled version of OpenAI's Whisper model for English speech recognition, offering 6x faster inference, 49% fewer par…
274112active
tmoroney/auto-subs
AutoSubs is a local-first desktop application that generates AI subtitles on-device using Whisper, Moonshine, and Parakeet models, with spe…
944091active
umlx5h/LLPlayer
LLPlayer is a Windows media player built for language learning, featuring dual subtitles, AI-generated subtitles via Whisper ASR, real-time…
784049active
MOSS-TTS
MOSS-TTS-Nano is an open-source 0.1B-parameter multilingual speech generation (TTS) model from MOSI.AI and the OpenMOSS team, designed for …
574031active
hanshuaikang/AI-Media2Doc
A self-hostable web application that uses AI large language models to convert video and audio into various document styles such as Xiaohong…
573994active
QwenLM/Qwen3-Omni
Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an…
523980active
mushan0x0/AI0x0.com
AI 0x0 is a multimodal, multi-model desktop AI assistant that lives as a floating ball on macOS and Windows, letting users invoke AI querie…
343938active
digimata/quill
Quill is an ultra-minimalist, fully local macOS meeting recorder and transcriber that runs as a single Swift menu-bar binary. It records mi…
553874active
Nekogram/Nekogram
Nekogram is an open-source third-party Telegram client for Android, forked from the official Telegram for Android, adding useful modificati…
933831active
HumanAIGC-Engineering/OpenAvatarChat
OpenAvatarChat is a modular interactive digital human (talking avatar) chat application that combines ASR, LLM, TTS, and avatar rendering c…
743723active
f/textream
Textream is a free, open-source teleprompter app for Mac, iPhone, and iPad that highlights your script in real time as you speak, using wor…
823686active
murtaza-nasir/speakr
Speakr is a self-hosted web application that transcribes audio recordings and turns them into organized, searchable notes. It supports mult…
843676active
sukeesh/Jarvis
Jarvis is a command-line personal assistant for Linux, macOS, and Windows written in Python. It offers 15+ task categories including weathe…
483645active
kaldi-asr/kaldi
Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea…
5215469maintenance
Melvin-Abraham/Google-Assistant-Unofficial-Desktop-Client
An unofficial cross-platform desktop client for Google Assistant built on the Google Assistant SDK using Electron. It provides a Chrome OS-…
233631active
speaches-ai/speaches
Speaches is an OpenAI API-compatible self-hosted server for speech-to-text (via faster-whisper), translation, and text-to-speech (via Kokor…
783621active
SakiRinn/LiveCaptions-Translator
A lightweight Windows application that combines the built-in Windows 11 LiveCaptions speech-to-text feature with translation APIs (includin…
773585active
Soul-AILab/SoulX-Podcast
SoulX-Podcast is the official inference codebase for a text-to-speech model that generates long-form, multi-turn, multi-speaker podcast-sty…
423535active
common-voice/common-voice
Mozilla Common Voice is a web platform for crowdsourcing voice donations to build public-domain speech datasets for training voice recognit…
953484active
antiboredom/videogrep
Videogrep is a Python command line tool that searches through dialog in video or audio files using subtitle tracks or speech transcriptions…
233461active
n3d1117/chatgpt-telegram-bot
A self-hosted Telegram bot written in Python that integrates with OpenAI's official ChatGPT, DALL·E, and Whisper APIs to answer questions, …
333460active
Kedreamix/Linly-Talker
Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, …
483436active
Qwen3-ASR
Qwen3-ASR is a family of open-source speech recognition models from Alibaba's Qwen team, supporting ASR and language identification across …
553423active
WEIFENG2333/AsrTools
AsrTools is a Python desktop application with a PyQt5-based GUI that converts audio and video files to text using online ASR engines, witho…
483423active
AnySoftKeyboard/AnySoftKeyboard
AnySoftKeyboard is a free, open-source on-screen keyboard application for Android supporting 50+ languages via external language packs. It …
803367active
xenova/whisper-web
A browser-based speech recognition app that runs OpenAI's Whisper models entirely client-side using Transformers.js. It transcribes audio w…
303338active
Kedreamix/Linly-Dubbing
Linly-Dubbing is an intelligent multi-language AI dubbing and video translation tool that combines speech recognition (WhisperX, FunASR), L…
273331active
ahmetoner/whisper-asr-webservice
A Dockerized REST webservice that wraps OpenAI Whisper (plus Faster Whisper and WhisperX engines) for automatic speech recognition. It expo…
933326active
Open-Less/openless
OpenLess is an open-source cross-platform voice input application that lets users hold a hotkey, speak, and have AI-polished text inserted …
813321active
XiaoMi/xiaomi-miloco
Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an…
823292active
timerring/bilive
BILIVE is a Python application that records Bilibili live streams and danmaku 24/7, then automatically renders danmaku and AI-generated sub…
613275active
echo-loop/Echo-Loop
Echo Loop is an open-source Flutter-based English listening and speaking training app that guides learners through a structured listen-to-s…
803242active
VOICEVOX
VOICEVOX is a free, mid-quality text-to-speech (and singing synthesis) software whose editor is built with Electron, TypeScript, and Vue. I…
933231active
intel/acat
ACAT is an open-source assistive communication platform from Intel Labs, originally developed for Stephen Hawking, that helps people with r…
653230active
zai-org/GLM-4-Voice
GLM-4-Voice is an end-to-end bilingual (Chinese/English) speech dialogue model from Zhipu AI, built on GLM-4-9B with a speech tokenizer and…
223222active
wendy7756/AI-Video-Transcriber
An open-source AI tool that transcribes, summarizes, and archives videos and podcasts from 30+ platforms (YouTube, TikTok, Bilibili, etc.) …
623215active
Purfview/whisper-standalone-win
Standalone Windows/Linux/macOS executables of OpenAI's Whisper and Faster-Whisper that transcribe audio and video to text and subtitles wit…
503163active
BayLing-Models/BayLing-Speech
LLaMA-Omni is an end-to-end speech interaction model built on Llama-3.1-8B-Instruct that generates simultaneous text and speech responses f…
323146active
matthartman/ghost-pepper
A free, open-source macOS menu bar app providing fully on-device speech-to-text dictation and meeting transcription using local Whisper, Pa…
783140active
chenyme/Chenyme-AAVT
Chenyme-AAVT is a fully automated audio/video translation application that uses Whisper (faster-whisper) for speech recognition, large lang…
103128active
SuperCmdLabs/SuperCmd
SuperCmd is an open-source macOS launcher combining Raycast-compatible extensions, hold-to-speak dictation, text-to-speech, AI chat and age…
753124active
modelscope/3D-Speaker
3D-Speaker is an open-source Python toolkit for single- and multi-modal speaker verification, speaker recognition, and speaker diarization,…
563121active
HeyWillow/willow
Willow is an open source, self-hosted voice assistant platform for ESP32-S3-BOX hardware, designed as a privacy-focused alternative to Amaz…
863097active
theajack/cnchar
cnchar is a comprehensive TypeScript library for Chinese character processing, offering pinyin conversion, stroke counts, stroke order draw…
663083active
futo-org/android-keyboard
FUTO Keyboard is a privacy-focused Android keyboard app forked from AOSP's LatinIME, offering offline voice input, swipe typing, autocorrec…
863081active
elevenlabs/elevenlabs-python
The official Python SDK for the ElevenLabs API, providing programmatic access to text-to-speech, speech-to-text, voice cloning, dubbing, mu…
933078active
sonos/tract
Tract is Sonos' tiny, self-contained neural-network inference engine written in Rust. It loads ONNX, TensorFlow/TFLite, and NNEF models, op…
993045active
kyutai-labs/delayed-streams-modeling
Kyutai's repository of Speech-To-Text and Text-To-Speech models built on the Delayed Streams Modeling framework, with implementations in Py…
473017active
off-grid-ai/OGAM
Off Grid AI (OGAM) is a cross-platform mobile and desktop application that runs AI entirely on-device: GGUF LLM chat with vision, Whisper s…
783002active
CheshireCC/faster-whisper-GUI
A desktop GUI application built with PySide6 for running faster-whisper and whisperX speech-to-text transcription. It lets users transcribe…
712991active
KevinWang676/Bark-Voice-Cloning
A one-click hub of Gradio Web UIs and Colab notebooks for open-source voice cloning, TTS, and voice conversion models including Bark, GPT-S…
672947active
davabase/whisper_real_time
A Python demo application that performs real-time speech-to-text transcription using OpenAI's Whisper model. It records audio continuously …
392941active
facebookresearch/omnilingual-asr
An open-source multilingual speech recognition library from Meta AI supporting over 1,600 languages, including hundreds never previously co…
522898active
OpenMind/OM1
OM1 is a modular AI runtime and hardware abstraction layer for building multimodal AI agents that run on physical robots and in simulators.…
822897active
kitlangton/Hex
Hex is a macOS application that converts your voice to text: press-and-hold a global hotkey to record, and it transcribes on-device and pas…
852890active
openai/openai-fm
OpenAI.fm is an interactive web demo showcasing OpenAI's text-to-speech models, built with Next.js and the OpenAI Speech API. It lets users…
512887active
thepersonalaicompany/amurex
Amurex is an open-source Chrome extension that acts as an AI meeting copilot for Google Meet and MS Teams. It provides real-time suggestion…
342868active

← prev page 2 / 9 next →