function: speech-recognition
801 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| whotto/Video_note_generator A Python tool that converts video URLs into polished Xiaohongshu (Little Red Book) notes and blog articles. It downloads videos, transcribe… | 43 | 1808 | active |
| jd-opensource/JoyAI-VL-Interaction JoyAI-VL-Interaction is an open 8B-scale vision-language interaction model with a complete deployable real-time streaming system, including… | 58 | 1806 | active |
| tuya/TuyaOpen TuyaOpen is an open-source, cross-platform C/C++ SDK and IoT OS for building AI-agent hardware on Tuya T-series MCUs, ESP32, Beken, LN882H,… | 87 | 1804 | active |
| skalesapp/skales Skales is a personal AI agent desktop and mobile application that runs locally on Windows, macOS, Linux, Android, and iOS, executing multi-… | 82 | 1727 | active |
| AutoArk/EVA-OS EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining … | 68 | 1712 | active |
| whitphx/streamlit-webrtc A Python library that adds real-time video and audio streaming to Streamlit apps via WebRTC. It lets developers process live camera/microph… | 94 | 1706 | active |
| software-mansion/react-native-executorch React Native ExecuTorch is a declarative React Native library for running AI models on-device, powered by Meta's ExecuTorch runtime. It shi… | 88 | 1702 | active |
| stack-chan/stack-chan Stack-chan is an open-source, JavaScript/TypeScript-driven desktop robot built on M5Stack (CoreS3) hardware, with firmware, MOD application… | 95 | 1674 | active |
| elixir-nx/bumblebee Bumblebee is an Elixir library providing pre-trained neural network models built on Axon, with integration for downloading models from Hugg… | 89 | 1662 | active |
| LlamaEdge/LlamaEdge LlamaEdge is a lightweight Rust and WasmEdge-based runtime for running open-source LLMs locally or on edge devices, with CLI chat apps and … | 65 | 1650 | active |
| fastclaw-ai/weclaw WeClaw is a Go-based bridge that connects WeChat to AI coding agents like Claude, Codex, Gemini, and Kimi via ACP, CLI, or OpenAI-compatibl… | 63 | 1644 | active |
| ZiqiaoPeng/SyncTalk SyncTalk is the official PyTorch implementation of a CVPR 2024 paper that synthesizes speech-driven, synchronized talking head videos using… | 46 | 1626 | active |
| yincongcyincong/MuseBot MuseBot is a self-hosted Go chatbot application that connects messaging platforms (Telegram, Discord, Slack, Lark/Feishu, DingTalk, WeCom, … | 78 | 1623 | active |
| LYiHub/mad-professor-public A desktop AI companion application for reading academic papers, built with PyQt6. It combines PDF parsing and translation, RAG-based retrie… | 27 | 1604 | active |
| waybarrios/vllm-mlx vllm-mlx is a vLLM-style LLM inference server for Apple Silicon Macs built on native MLX, exposing both OpenAI and Anthropic compatible API… | 82 | 1547 | active |
| RTGS2017/NagaAgent NagaAgent is a desktop AI personal assistant framework with an anime (Live2D) virtual character, built in Python with an Electron front end… | 83 | 1544 | active |
| LakshmanTurlapati/Review-Gate Review-Gate is a rule and MCP tool for the Cursor IDE that keeps the AI agent in an interactive loop, waiting for follow-up text, voice, or… | 68 | 1528 | active |
| tin2tin/Pallaidium Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE… | 75 | 1520 | active |
| met4citizen/TalkingHead A JavaScript library for real-time lip-synced talking avatars, rendering full-body 3D characters in the browser. It supports text-to-speech… | 67 | 1505 | active |
| microsoft/ai-dev-gallery AI Dev Gallery is a Windows application from Microsoft that lets developers explore over 25 interactive samples powered by local AI models … | 66 | 1497 | active |
| fynnfluegge/rocketnotes Rocketnotes is a web-based Markdown note-taking app with integrated AI features such as chat with your documents, text completion, voice-to… | 69 | 1491 | active |
| import-ai/omnibox OmniBox is a cross-platform, self-hostable AI knowledge hub that lets users collect webpages, files, and voice notes, then parse, index, an… | 87 | 1481 | active |
| siddsachar/row-bot Row-Bot is a local-first desktop AI assistant and workbench that combines chat, durable memory, a personal knowledge graph, tool use, paren… | 81 | 1455 | active |
| qwersyk/Newelle Newelle is a GTK4/GNOME virtual assistant application for Linux that connects to local (Ollama, Llama.cpp) and cloud LLM providers. It supp… | 96 | 1449 | active |
| watson-developer-cloud/python-sdk The official Python client library (pip package ibm-watson) for accessing IBM Watson AI services such as language, speech, and vision APIs.… | 68 | 1449 | active |
| v-modal/vmodal_sdk_flutter VModal for Flutter is a Dart/Flutter SDK that adds multimodal video search (semantic, ASR, and OCR based) and streamed video uploads to And… | 57 | 1448 | active |
| ByteDance-Seed/m3-agent M3-Agent is a multimodal agent framework from ByteDance Seed that processes real-time visual and auditory inputs to build entity-centric lo… | 48 | 1445 | active |
| voice-cloning-app/Voice-Cloning-App A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom… | 23 | 1440 | active |
| team-reflect/reflect-open An open-source, local-first note-taking app for Mac and iPhone that stores notes as plain Markdown files with daily notes, wiki links, back… | 80 | 1438 | active |
| ammaarreshi/gemma-chat An open-source Electron app that runs Google's Gemma 4 models locally on Apple Silicon via Apple's MLX framework, providing a chat interfac… | 50 | 1424 | active |
| Zejun-Yang/AniPortrait AniPortrait is a Python framework from Tencent that generates photorealistic portrait animations from an audio clip and a reference image, … | 25 | 5021 | maintenance |
| ARahim3/mlx-tune A Python library for fine-tuning LLMs, vision-language, audio (TTS/STT), embedding, OCR, and JEPA models natively on Apple Silicon Macs usi… | 75 | 1389 | active |
| moyangzhan/langchain4j-aideepin LangChain4j-AIDeepin is an open-source, self-hostable AI application platform built on Spring Boot, LangChain4j, and LangGraph4j with Vue 3… | 91 | 1361 | active |
| MoonInTheRiver/DiffSinger Official PyTorch implementation of DiffSinger, an AAAI 2022 paper on singing voice synthesis and text-to-speech using a shallow diffusion m… | 65 | 4851 | maintenance |
| ParisNeo/lollms-webui LoLLMs WebUI is a local, single-user web interface for running large language models and multimodal AI systems, supporting hundreds of mode… | 71 | 4785 | maintenance |
| morettt/my-neuro An open-source AI desktop companion framework inspired by Neuro-sama, letting users build a customizable Live2D character with sub-second v… | 84 | 1342 | active |
| joey-zhou/xiaozhi-esp32-server-java A Java enterprise-grade server and management platform for the Xiaozhi ESP32 AI voice assistant hardware, providing a full front-end/back-e… | 69 | 1336 | active |
| jishengpeng/WavTokenizer WavTokenizer is a state-of-the-art discrete neural audio codec that compresses speech, music, and general audio into only 40 or 75 discrete… | 27 | 1316 | active |
| Capsize-Games/airunner AI Runner is a privacy-focused desktop application for running local AI models offline, combining an AI chat companion with voice conversat… | 79 | 1314 | active |
| yeyupiaoling/VoiceprintRecognition-Pytorch A PyTorch-based voiceprint recognition (speaker recognition) framework implementing models such as ECAPA-TDNN, ResNetSE, ERes2Net, and CAM+… | 58 | 1312 | active |
| ttop32/MouseTooltipTranslator A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports … | 95 | 1300 | active |
| HG-ha/MTools MTools is a cross-platform desktop application built with Python and Flet that bundles image processing, audio/video editing, text operatio… | 83 | 1298 | active |
| wisupai/e2m E2M is a Python library that parses and converts many file types (doc, docx, epub, html, url, pdf, ppt, pptx, mp3, m4a) into Markdown using… | 23 | 1294 | active |
| Phantom-video/HuMo HuMo is a research model and Python codebase from Tsinghua University and ByteDance for human-centric video generation using collaborative … | 46 | 1283 | active |
| PrunaAI/pruna Pruna is an open-source Python model optimization framework that makes AI models faster, smaller, cheaper, and greener via caching, quantiz… | 84 | 1275 | active |
| jofizcd/Soul-of-Waifu Soul of Waifu is an open-source desktop application for creating AI companion characters with Live2D/VRM avatars, voice chat, and persisten… | 94 | 1263 | active |
| myreader-io/myGPTReader myGPTReader is a Slack bot powered by ChatGPT that reads and summarizes webpages, documents (eBooks, PDF, DOCX), and YouTube videos, and su… | 60 | 4419 | maintenance |
| lcoutodemos/clui-cc Clui CC is a macOS desktop overlay application that wraps the Claude Code CLI in a floating, transparent pill interface with multi-tab sess… | 48 | 1225 | active |
| kellyvv/PhoneClaw PhoneClaw is a mobile-native local AI agent framework that turns phones into on-device agent runtimes, running Gemma models via LiteRT and … | 77 | 1220 | active |
| GML-MMGroup/GMTalker GMTalker is an interactive 3D digital human system rendered with Unreal Engine, integrating speech recognition, speech synthesis, natural l… | 46 | 1217 | active |
| metavoiceio/metavoice-src MetaVoice-1B is a 1.2B parameter foundational text-to-speech model trained on 100K hours of speech, focused on emotional rhythm and tone in… | 26 | 4205 | maintenance |
| qualcomm/ai-hub-models Qualcomm AI Hub Models is a curated collection of 300+ state-of-the-art machine learning models (vision, audio, speech, generative AI) pre-… | 89 | 1195 | active |
| warmshao/FasterLivePortrait A real-time portrait animation application based on LivePortrait that animates still photos or videos using a driving video, image, audio, … | 38 | 1174 | active |
| keinsaasforever/better-chatbot Keinsaas Navigator (formerly Better Chatbot) is an open-source, self-hostable AI chatbot workspace built with Next.js and the Vercel AI SDK… | 71 | 1169 | active |
| whyiyhw/chatgpt-wechat A self-hosted Go application that lets users safely use LLM assistants (ChatGPT, Gemini, DeepSeek, Dify workflows) inside WeChat by relayin… | 59 | 1169 | active |
| glink25/Cent Cent is a free, open-source collaborative bookkeeping (expense tracking) web app built as a pure-frontend PWA that stores ledger data as JS… | 62 | 1168 | active |
| StreamerHelper/web-server The backend service for StreamerHelper, a self-hosted livestream recording system. It polls live status from platforms like Bilibili, Huya,… | 76 | 1153 | active |
| Woolverine94/biniou biniou is a self-hosted web UI for 30+ generative AI models covering image, video, audio, and text generation, built with Gradio and Huggin… | 63 | 1150 | active |
| cloudflare/ai A monorepo of TypeScript packages and examples for building AI-powered applications on Cloudflare. It provides Vercel AI SDK and TanStack A… | 80 | 1148 | active |
| smthemex/ComfyUI_Sonic A ComfyUI custom node implementing the Sonic method for audio-driven portrait animation, generating talking-head videos from a single portr… | 56 | 1140 | active |
| laravel/ai The Laravel AI SDK is a PHP package offering a unified, expressive API for interacting with AI providers such as OpenAI, Anthropic, and Gem… | 84 | 1139 | active |
| HITsz-TMG/Uni-MoE Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,… | 68 | 1116 | active |
| ATH-MaaS/Pixelle-MCP Pixelle MCP is an open-source omnimodal AIGC framework that converts ComfyUI workflows (local or RunningHub cloud) into MCP tools with zero… | 44 | 1104 | active |
| yerfor/Real3DPortrait Official PyTorch implementation of Real3D-Portrait, an ICLR 2024 Spotlight paper for one-shot realistic 3D talking portrait synthesis. It g… | 26 | 1091 | active |
| WhiskeyCoder/Qwen3-Audiobook-Converter A Python CLI tool that converts documents (PDF, EPUB, DOCX, DOC, TXT) into audiobooks using the Qwen3 TTS voice model running locally via a… | 49 | 1081 | active |
| memoavatar/memo MEMO is an open-weight diffusion model for generating expressive, identity-consistent talking videos from a single reference image and an a… | 40 | 1070 | active |
| X-LANCE/SLAM-LLM SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc… | 55 | 1056 | active |
| Dicklesworthstone/swiss_army_llama A FastAPI-based REST service that exposes local LLM capabilities including text embeddings, completions, semantic similarity, and semantic … | 32 | 1056 | active |
| tegnike/aituber-kit AITuberKit is an all-in-one web application toolkit for building and deploying AI character chat experiences, including streaming-oriented … | 89 | 1054 | active |
| agents-flex/agents-flex Agents-Flex is a lightweight, modular Java framework for building AI applications and agents, positioned as a Java counterpart to Spring AI… | 93 | 1046 | active |
| k2-fsa/ZipVoice ZipVoice is a series of fast, high-quality zero-shot text-to-speech models based on flow matching, with a compact 123M-parameter Zipformer-… | 43 | 1045 | active |
| ccrma/chuck ChucK is an open-source, strongly-timed programming language for real-time sound synthesis and music creation, with a unique time-based con… | 73 | 1038 | active |
| Prajwal100/Complete-Ecommerce-in-laravel-10 A full-featured e-commerce website built with Laravel 10, including a storefront with cart, wishlist, and order tracking, plus an admin das… | 61 | 1015 | active |
| jianjieyiban/JJYB_AI_VideoAutoCut JJYB_AI 智剪 is a local-first desktop AI video creation workbench that combines material analysis, smart shot segmentation, commentary script… | 70 | 1013 | active |
| Soul-AILab/SoulX-FlashHead SoulX-FlashHead is a 1.3B-parameter framework for high-fidelity, infinite-length, real-time streaming talking-head portrait video generatio… | 53 | 1011 | active |
| PlayVoice/whisper-vits-svc A PyTorch-based singing voice conversion and voice cloning engine built on VITS with Whisper, BigVGAN, and diffusion components. It lets us… | 23 | 2864 | maintenance |
| weaigc/bingo Bingo is a self-hostable web application that recreates the New Bing (Bing AI / Copilot) chat interface, allowing access to Bing AI feature… | 28 | 2832 | maintenance |
| SCUTlihaoyu/open-chat-video-editor An open-source Python tool that automatically generates short videos from a short text prompt or a web URL, producing narration, background… | 29 | 2813 | maintenance |
| yerfor/GeneFace GeneFace is the official PyTorch implementation of an ICLR 2023 paper on generalized, high-fidelity audio-driven 3D talking face synthesis … | 21 | 2657 | maintenance |
| waooAI/waoowaoo waoowaoo is a self-hosted AI film and video production studio that turns novel or script text into complete videos, automatically generatin… | 74 | 13835 | experimental |
| SUSI.AI SUSI.AI is an open-source personal assistant platform whose Java server holds the assistant's 'intelligence', answering chat and voice quer… | 10 | 2521 | maintenance |
| r9y9/wavenet_vocoder A PyTorch implementation of the WaveNet vocoder that generates high-quality raw speech waveforms conditioned on acoustic features like mel-… | 23 | 2376 | maintenance |
| circlestarzero/EX-chatGPT Ex-ChatGPT is a Python web application that lets ChatGPT call external APIs (Google, WolframAlpha, WikiMedia) to give more accurate, up-to-… | 10 | 1959 | maintenance |
| johnwheeler/flask-ask Flask-Ask is a Flask extension that simplifies building Alexa Skills for Amazon Echo devices in Python. It maps Alexa intents to view funct… | 32 | 1909 | maintenance |
| OkGoDoIt/OpenAI-API-dotnet An unofficial C#/.NET SDK wrapping the OpenAI API, covering chat completions (GPT-3.5/4), DALL-E image generation, embeddings, moderation, … | 23 | 1893 | maintenance |
| Kyubyong/tacotron A heavily documented TensorFlow implementation of Tacotron, a fully end-to-end text-to-speech synthesis model. It includes training, prepro… | 32 | 1832 | maintenance |
| microsoft/i-Code Microsoft's i-Code is a collection of research models and frameworks for integrative, composable multimodal AI spanning vision, language, a… | 32 | 1703 | maintenance |
| kan-bayashi/ParallelWaveGAN Unofficial PyTorch implementations of non-autoregressive neural vocoders including Parallel WaveGAN, MelGAN, Multi-band MelGAN, HiFi-GAN, a… | 23 | 1645 | maintenance |
| Delta-ML/delta DELTA is a deep learning based end-to-end natural language and speech processing platform built on TensorFlow and Python 3. It provides one… | 10 | 1607 | maintenance |
| HumanAIGC/EMO EMO (Emote Portrait Alive) is a research codebase from Alibaba's Institute for Intelligent Computing that generates expressive talking port… | 25 | 7594 | experimental |
| vercel/modelfusion ModelFusion is a TypeScript library that provides a unified, vendor-neutral abstraction layer for integrating AI models into JavaScript and… | 10 | 1319 | maintenance |
| YuanxunLu/LiveSpeechPortraits A PyTorch implementation of the SIGGRAPH Asia 2021 paper 'Live Speech Portraits', which generates photorealistic personalized talking-head … | 32 | 1283 | maintenance |
| ARM-software/ML-KWS-for-MCU TensorFlow models and training scripts for keyword spotting (wake-word detection) on Arm Cortex-M microcontrollers, accompanying the 'Hello… | 32 | 1249 | maintenance |
| mravanelli/SincNet SincNet is a PyTorch neural architecture that processes raw audio waveforms using parametrized sinc band-pass filters in the first convolut… | 32 | 1243 | maintenance |
| clovaai/voxceleb_trainer A PyTorch framework for training and evaluating speaker recognition and verification models on the VoxCeleb datasets. It implements multipl… | 67 | 1175 | maintenance |
| IliasHad/edit-mind Edit Mind is a local-first video knowledge base that indexes video libraries with multi-modal AI analysis (Whisper transcription, YOLO obje… | 76 | 1789 | experimental |
| HumanMLLM/R1-Omni R1-Omni is a research project applying Reinforcement Learning with Verifiable Reward (RLVR) to an omni-multimodal large language model for … | 26 | 1022 | experimental |
| openai/openai-realtime-api-beta A Node.js and browser reference client library for OpenAI's Realtime API, enabling real-time voice and text conversations with GPT models. … | 22 | 1016 | experimental |
| NVIDIA/ChatRTX ChatRTX is a Windows demo application for building personalized RAG chatbots on local RTX GPUs using TensorRT-LLM, NVIDIA NIM, and LlamaInd… | 10 | 3121 | abandoned |
| plamoni/SiriProxy SiriProxy is a Ruby-based tampering proxy server for Apple's Siri assistant that intercepts Siri traffic and lets developers write custom p… | 10 | 2118 | abandoned |