Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: speech-recognition

801 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
whotto/Video_note_generator
A Python tool that converts video URLs into polished Xiaohongshu (Little Red Book) notes and blog articles. It downloads videos, transcribe…
431808active
jd-opensource/JoyAI-VL-Interaction
JoyAI-VL-Interaction is an open 8B-scale vision-language interaction model with a complete deployable real-time streaming system, including…
581806active
tuya/TuyaOpen
TuyaOpen is an open-source, cross-platform C/C++ SDK and IoT OS for building AI-agent hardware on Tuya T-series MCUs, ESP32, Beken, LN882H,…
871804active
skalesapp/skales
Skales is a personal AI agent desktop and mobile application that runs locally on Windows, macOS, Linux, Android, and iOS, executing multi-…
821727active
AutoArk/EVA-OS
EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining …
681712active
whitphx/streamlit-webrtc
A Python library that adds real-time video and audio streaming to Streamlit apps via WebRTC. It lets developers process live camera/microph…
941706active
software-mansion/react-native-executorch
React Native ExecuTorch is a declarative React Native library for running AI models on-device, powered by Meta's ExecuTorch runtime. It shi…
881702active
stack-chan/stack-chan
Stack-chan is an open-source, JavaScript/TypeScript-driven desktop robot built on M5Stack (CoreS3) hardware, with firmware, MOD application…
951674active
elixir-nx/bumblebee
Bumblebee is an Elixir library providing pre-trained neural network models built on Axon, with integration for downloading models from Hugg…
891662active
LlamaEdge/LlamaEdge
LlamaEdge is a lightweight Rust and WasmEdge-based runtime for running open-source LLMs locally or on edge devices, with CLI chat apps and …
651650active
fastclaw-ai/weclaw
WeClaw is a Go-based bridge that connects WeChat to AI coding agents like Claude, Codex, Gemini, and Kimi via ACP, CLI, or OpenAI-compatibl…
631644active
ZiqiaoPeng/SyncTalk
SyncTalk is the official PyTorch implementation of a CVPR 2024 paper that synthesizes speech-driven, synchronized talking head videos using…
461626active
yincongcyincong/MuseBot
MuseBot is a self-hosted Go chatbot application that connects messaging platforms (Telegram, Discord, Slack, Lark/Feishu, DingTalk, WeCom, …
781623active
LYiHub/mad-professor-public
A desktop AI companion application for reading academic papers, built with PyQt6. It combines PDF parsing and translation, RAG-based retrie…
271604active
waybarrios/vllm-mlx
vllm-mlx is a vLLM-style LLM inference server for Apple Silicon Macs built on native MLX, exposing both OpenAI and Anthropic compatible API…
821547active
RTGS2017/NagaAgent
NagaAgent is a desktop AI personal assistant framework with an anime (Live2D) virtual character, built in Python with an Electron front end…
831544active
LakshmanTurlapati/Review-Gate
Review-Gate is a rule and MCP tool for the Cursor IDE that keeps the AI agent in an interactive loop, waiting for follow-up text, voice, or…
681528active
tin2tin/Pallaidium
Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE…
751520active
met4citizen/TalkingHead
A JavaScript library for real-time lip-synced talking avatars, rendering full-body 3D characters in the browser. It supports text-to-speech…
671505active
microsoft/ai-dev-gallery
AI Dev Gallery is a Windows application from Microsoft that lets developers explore over 25 interactive samples powered by local AI models …
661497active
fynnfluegge/rocketnotes
Rocketnotes is a web-based Markdown note-taking app with integrated AI features such as chat with your documents, text completion, voice-to…
691491active
import-ai/omnibox
OmniBox is a cross-platform, self-hostable AI knowledge hub that lets users collect webpages, files, and voice notes, then parse, index, an…
871481active
siddsachar/row-bot
Row-Bot is a local-first desktop AI assistant and workbench that combines chat, durable memory, a personal knowledge graph, tool use, paren…
811455active
qwersyk/Newelle
Newelle is a GTK4/GNOME virtual assistant application for Linux that connects to local (Ollama, Llama.cpp) and cloud LLM providers. It supp…
961449active
watson-developer-cloud/python-sdk
The official Python client library (pip package ibm-watson) for accessing IBM Watson AI services such as language, speech, and vision APIs.…
681449active
v-modal/vmodal_sdk_flutter
VModal for Flutter is a Dart/Flutter SDK that adds multimodal video search (semantic, ASR, and OCR based) and streamed video uploads to And…
571448active
ByteDance-Seed/m3-agent
M3-Agent is a multimodal agent framework from ByteDance Seed that processes real-time visual and auditory inputs to build entity-centric lo…
481445active
voice-cloning-app/Voice-Cloning-App
A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom…
231440active
team-reflect/reflect-open
An open-source, local-first note-taking app for Mac and iPhone that stores notes as plain Markdown files with daily notes, wiki links, back…
801438active
ammaarreshi/gemma-chat
An open-source Electron app that runs Google's Gemma 4 models locally on Apple Silicon via Apple's MLX framework, providing a chat interfac…
501424active
Zejun-Yang/AniPortrait
AniPortrait is a Python framework from Tencent that generates photorealistic portrait animations from an audio clip and a reference image, …
255021maintenance
ARahim3/mlx-tune
A Python library for fine-tuning LLMs, vision-language, audio (TTS/STT), embedding, OCR, and JEPA models natively on Apple Silicon Macs usi…
751389active
moyangzhan/langchain4j-aideepin
LangChain4j-AIDeepin is an open-source, self-hostable AI application platform built on Spring Boot, LangChain4j, and LangGraph4j with Vue 3…
911361active
MoonInTheRiver/DiffSinger
Official PyTorch implementation of DiffSinger, an AAAI 2022 paper on singing voice synthesis and text-to-speech using a shallow diffusion m…
654851maintenance
ParisNeo/lollms-webui
LoLLMs WebUI is a local, single-user web interface for running large language models and multimodal AI systems, supporting hundreds of mode…
714785maintenance
morettt/my-neuro
An open-source AI desktop companion framework inspired by Neuro-sama, letting users build a customizable Live2D character with sub-second v…
841342active
joey-zhou/xiaozhi-esp32-server-java
A Java enterprise-grade server and management platform for the Xiaozhi ESP32 AI voice assistant hardware, providing a full front-end/back-e…
691336active
jishengpeng/WavTokenizer
WavTokenizer is a state-of-the-art discrete neural audio codec that compresses speech, music, and general audio into only 40 or 75 discrete…
271316active
Capsize-Games/airunner
AI Runner is a privacy-focused desktop application for running local AI models offline, combining an AI chat companion with voice conversat…
791314active
yeyupiaoling/VoiceprintRecognition-Pytorch
A PyTorch-based voiceprint recognition (speaker recognition) framework implementing models such as ECAPA-TDNN, ResNetSE, ERes2Net, and CAM+…
581312active
ttop32/MouseTooltipTranslator
A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports …
951300active
HG-ha/MTools
MTools is a cross-platform desktop application built with Python and Flet that bundles image processing, audio/video editing, text operatio…
831298active
wisupai/e2m
E2M is a Python library that parses and converts many file types (doc, docx, epub, html, url, pdf, ppt, pptx, mp3, m4a) into Markdown using…
231294active
Phantom-video/HuMo
HuMo is a research model and Python codebase from Tsinghua University and ByteDance for human-centric video generation using collaborative …
461283active
PrunaAI/pruna
Pruna is an open-source Python model optimization framework that makes AI models faster, smaller, cheaper, and greener via caching, quantiz…
841275active
jofizcd/Soul-of-Waifu
Soul of Waifu is an open-source desktop application for creating AI companion characters with Live2D/VRM avatars, voice chat, and persisten…
941263active
myreader-io/myGPTReader
myGPTReader is a Slack bot powered by ChatGPT that reads and summarizes webpages, documents (eBooks, PDF, DOCX), and YouTube videos, and su…
604419maintenance
lcoutodemos/clui-cc
Clui CC is a macOS desktop overlay application that wraps the Claude Code CLI in a floating, transparent pill interface with multi-tab sess…
481225active
kellyvv/PhoneClaw
PhoneClaw is a mobile-native local AI agent framework that turns phones into on-device agent runtimes, running Gemma models via LiteRT and …
771220active
GML-MMGroup/GMTalker
GMTalker is an interactive 3D digital human system rendered with Unreal Engine, integrating speech recognition, speech synthesis, natural l…
461217active
metavoiceio/metavoice-src
MetaVoice-1B is a 1.2B parameter foundational text-to-speech model trained on 100K hours of speech, focused on emotional rhythm and tone in…
264205maintenance
qualcomm/ai-hub-models
Qualcomm AI Hub Models is a curated collection of 300+ state-of-the-art machine learning models (vision, audio, speech, generative AI) pre-…
891195active
warmshao/FasterLivePortrait
A real-time portrait animation application based on LivePortrait that animates still photos or videos using a driving video, image, audio, …
381174active
keinsaasforever/better-chatbot
Keinsaas Navigator (formerly Better Chatbot) is an open-source, self-hostable AI chatbot workspace built with Next.js and the Vercel AI SDK…
711169active
whyiyhw/chatgpt-wechat
A self-hosted Go application that lets users safely use LLM assistants (ChatGPT, Gemini, DeepSeek, Dify workflows) inside WeChat by relayin…
591169active
glink25/Cent
Cent is a free, open-source collaborative bookkeeping (expense tracking) web app built as a pure-frontend PWA that stores ledger data as JS…
621168active
StreamerHelper/web-server
The backend service for StreamerHelper, a self-hosted livestream recording system. It polls live status from platforms like Bilibili, Huya,…
761153active
Woolverine94/biniou
biniou is a self-hosted web UI for 30+ generative AI models covering image, video, audio, and text generation, built with Gradio and Huggin…
631150active
cloudflare/ai
A monorepo of TypeScript packages and examples for building AI-powered applications on Cloudflare. It provides Vercel AI SDK and TanStack A…
801148active
smthemex/ComfyUI_Sonic
A ComfyUI custom node implementing the Sonic method for audio-driven portrait animation, generating talking-head videos from a single portr…
561140active
laravel/ai
The Laravel AI SDK is a PHP package offering a unified, expressive API for interacting with AI providers such as OpenAI, Anthropic, and Gem…
841139active
HITsz-TMG/Uni-MoE
Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,…
681116active
ATH-MaaS/Pixelle-MCP
Pixelle MCP is an open-source omnimodal AIGC framework that converts ComfyUI workflows (local or RunningHub cloud) into MCP tools with zero…
441104active
yerfor/Real3DPortrait
Official PyTorch implementation of Real3D-Portrait, an ICLR 2024 Spotlight paper for one-shot realistic 3D talking portrait synthesis. It g…
261091active
WhiskeyCoder/Qwen3-Audiobook-Converter
A Python CLI tool that converts documents (PDF, EPUB, DOCX, DOC, TXT) into audiobooks using the Qwen3 TTS voice model running locally via a…
491081active
memoavatar/memo
MEMO is an open-weight diffusion model for generating expressive, identity-consistent talking videos from a single reference image and an a…
401070active
X-LANCE/SLAM-LLM
SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc…
551056active
Dicklesworthstone/swiss_army_llama
A FastAPI-based REST service that exposes local LLM capabilities including text embeddings, completions, semantic similarity, and semantic …
321056active
tegnike/aituber-kit
AITuberKit is an all-in-one web application toolkit for building and deploying AI character chat experiences, including streaming-oriented …
891054active
agents-flex/agents-flex
Agents-Flex is a lightweight, modular Java framework for building AI applications and agents, positioned as a Java counterpart to Spring AI…
931046active
k2-fsa/ZipVoice
ZipVoice is a series of fast, high-quality zero-shot text-to-speech models based on flow matching, with a compact 123M-parameter Zipformer-…
431045active
ccrma/chuck
ChucK is an open-source, strongly-timed programming language for real-time sound synthesis and music creation, with a unique time-based con…
731038active
Prajwal100/Complete-Ecommerce-in-laravel-10
A full-featured e-commerce website built with Laravel 10, including a storefront with cart, wishlist, and order tracking, plus an admin das…
611015active
jianjieyiban/JJYB_AI_VideoAutoCut
JJYB_AI 智剪 is a local-first desktop AI video creation workbench that combines material analysis, smart shot segmentation, commentary script…
701013active
Soul-AILab/SoulX-FlashHead
SoulX-FlashHead is a 1.3B-parameter framework for high-fidelity, infinite-length, real-time streaming talking-head portrait video generatio…
531011active
PlayVoice/whisper-vits-svc
A PyTorch-based singing voice conversion and voice cloning engine built on VITS with Whisper, BigVGAN, and diffusion components. It lets us…
232864maintenance
weaigc/bingo
Bingo is a self-hostable web application that recreates the New Bing (Bing AI / Copilot) chat interface, allowing access to Bing AI feature…
282832maintenance
SCUTlihaoyu/open-chat-video-editor
An open-source Python tool that automatically generates short videos from a short text prompt or a web URL, producing narration, background…
292813maintenance
yerfor/GeneFace
GeneFace is the official PyTorch implementation of an ICLR 2023 paper on generalized, high-fidelity audio-driven 3D talking face synthesis …
212657maintenance
waooAI/waoowaoo
waoowaoo is a self-hosted AI film and video production studio that turns novel or script text into complete videos, automatically generatin…
7413835experimental
SUSI.AI
SUSI.AI is an open-source personal assistant platform whose Java server holds the assistant's 'intelligence', answering chat and voice quer…
102521maintenance
r9y9/wavenet_vocoder
A PyTorch implementation of the WaveNet vocoder that generates high-quality raw speech waveforms conditioned on acoustic features like mel-…
232376maintenance
circlestarzero/EX-chatGPT
Ex-ChatGPT is a Python web application that lets ChatGPT call external APIs (Google, WolframAlpha, WikiMedia) to give more accurate, up-to-…
101959maintenance
johnwheeler/flask-ask
Flask-Ask is a Flask extension that simplifies building Alexa Skills for Amazon Echo devices in Python. It maps Alexa intents to view funct…
321909maintenance
OkGoDoIt/OpenAI-API-dotnet
An unofficial C#/.NET SDK wrapping the OpenAI API, covering chat completions (GPT-3.5/4), DALL-E image generation, embeddings, moderation, …
231893maintenance
Kyubyong/tacotron
A heavily documented TensorFlow implementation of Tacotron, a fully end-to-end text-to-speech synthesis model. It includes training, prepro…
321832maintenance
microsoft/i-Code
Microsoft's i-Code is a collection of research models and frameworks for integrative, composable multimodal AI spanning vision, language, a…
321703maintenance
kan-bayashi/ParallelWaveGAN
Unofficial PyTorch implementations of non-autoregressive neural vocoders including Parallel WaveGAN, MelGAN, Multi-band MelGAN, HiFi-GAN, a…
231645maintenance
Delta-ML/delta
DELTA is a deep learning based end-to-end natural language and speech processing platform built on TensorFlow and Python 3. It provides one…
101607maintenance
HumanAIGC/EMO
EMO (Emote Portrait Alive) is a research codebase from Alibaba's Institute for Intelligent Computing that generates expressive talking port…
257594experimental
vercel/modelfusion
ModelFusion is a TypeScript library that provides a unified, vendor-neutral abstraction layer for integrating AI models into JavaScript and…
101319maintenance
YuanxunLu/LiveSpeechPortraits
A PyTorch implementation of the SIGGRAPH Asia 2021 paper 'Live Speech Portraits', which generates photorealistic personalized talking-head …
321283maintenance
ARM-software/ML-KWS-for-MCU
TensorFlow models and training scripts for keyword spotting (wake-word detection) on Arm Cortex-M microcontrollers, accompanying the 'Hello…
321249maintenance
mravanelli/SincNet
SincNet is a PyTorch neural architecture that processes raw audio waveforms using parametrized sinc band-pass filters in the first convolut…
321243maintenance
clovaai/voxceleb_trainer
A PyTorch framework for training and evaluating speaker recognition and verification models on the VoxCeleb datasets. It implements multipl…
671175maintenance
IliasHad/edit-mind
Edit Mind is a local-first video knowledge base that indexes video libraries with multi-modal AI analysis (Whisper transcription, YOLO obje…
761789experimental
HumanMLLM/R1-Omni
R1-Omni is a research project applying Reinforcement Learning with Verifiable Reward (RLVR) to an omni-multimodal large language model for …
261022experimental
openai/openai-realtime-api-beta
A Node.js and browser reference client library for OpenAI's Realtime API, enabling real-time voice and text conversations with GPT models. …
221016experimental
NVIDIA/ChatRTX
ChatRTX is a Windows demo application for building personalized RAG chatbots on local RTX GPUs using TensorRT-LLM, NVIDIA NIM, and LlamaInd…
103121abandoned
plamoni/SiriProxy
SiriProxy is a Ruby-based tampering proxy server for Apple's Siri assistant that intercepts Siri traffic and lets developers write custom p…
102118abandoned

← prev page 8 / 9 next →