function: speech-recognition
801 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| huggingface/transformers.js Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) … | 91 | 16270 | active |
| alibaba/MNN MNN is a lightweight, high-performance deep learning inference engine developed by Alibaba, supporting LLMs, vision, audio, and multimodal … | 93 | 15973 | active |
| tensorflow/tfjs-models A collection of pre-trained machine learning models ported to TensorFlow.js, published as npm packages for use in JavaScript projects. Mode… | 63 | 14793 | active |
| pipecat-ai/pipecat Pipecat is an open-source Python framework (BSD-2) for building real-time voice and multimodal conversational AI agents. It orchestrates sp… | 89 | 14769 | active |
| palmier-io/palmier-pro Palmier Pro is an open-source, Swift-native multi-track video editor for macOS (Apple Silicon) with built-in generative AI for video, image… | 77 | 13964 | active |
| xszyou/Fay Fay is an open-source Python digital human framework that connects 2.5D/3D/mobile/web digital humans or OpenAI-compatible LLMs to business … | 75 | 13455 | active |
| codexu/note-gen NoteGen is an open-source, cross-platform, local-first Markdown note-taking app built with Tauri and Next.js that separates quick capture (… | 86 | 12689 | active |
| h2oai/h2ogpt h2oGPT is an Apache-2.0 open-source application for chatting with local, private LLMs and querying/summarizing your own documents (PDFs, Wo… | 10 | 11969 | active |
| linyqh/NarratoAI NarratoAI is an open-source, self-hosted AI-powered video commentary and automated editing tool. It uses large language models to generate … | 88 | 10879 | active |
| sashabaranov/go-openai An unofficial Go client library for the OpenAI API covering chat completions, the Responses API, embeddings, image generation, audio transc… | 98 | 10748 | active |
| OpenVINO OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su… | 95 | 10740 | stable |
| Baiyuetribe/paper2gui Paper2GUI (now branded 小白兔AI / Xiaobaitu AI) is a desktop AI toolbox application that packages 50+ AI research models into install-free GUI… | 23 | 10709 | active |
| open-mmlab/Amphion Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, s… | 51 | 10271 | active |
| peng-zhihui/ElectronBot ElectronBot is an open-source mini desktop robot inspired by EVE from WALL-E, featuring 6 degrees of freedom, USB communication, and a disp… | 34 | 9466 | active |
| coaidev/coai CoAI.Dev is a self-hostable, multi-tenant AI platform combining a ChatGPT-style chat UI with an enterprise-grade unified LLM gateway suppor… | 56 | 9295 | active |
| lipku/LiveTalking LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit… | 89 | 9238 | active |
| modelscope/modelscope ModelScope is a Python library and ecosystem built on the 'Model-as-a-Service' concept, providing unified APIs to download, run inference o… | 98 | 9111 | active |
| adithya-s-k/omniparse OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into… | 49 | 7815 | active |
| enricoros/big-AGI Big-AGI is an open-source AI workspace web application that lets users chat with many state-of-the-art LLM providers (OpenAI, Anthropic, Ge… | 93 | 7102 | active |
| PaddlePaddle/models PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda… | 23 | 6932 | active |
| facebookresearch/fairseq Fairseq is a PyTorch-based sequence modeling toolkit from Facebook AI Research for training custom models for translation, summarization, l… | 10 | 32231 | maintenance |
| goldendict/goldendict GoldenDict is a feature-rich desktop dictionary lookup program that supports many dictionary formats (StarDict, Babylon, Lingvo DSL, Dictd,… | 44 | 6649 | active |
| makepad/makepad Makepad is a creative software development platform for Rust that provides a cross-platform UI runtime with a live-editable design DSL, com… | 67 | 6580 | active |
| TMElyralab/MuseTalk MuseTalk is a real-time, high-fidelity lip-sync model that modifies a face region in video according to input audio via latent space inpain… | 45 | 6459 | active |
| vllm-project/vllm-omni vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer… | 83 | 6369 | active |
| PaddlePaddle/PaddleX PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti… | 92 | 6251 | active |
| JimmyLv/BibiGPT-v1 BibiGPT is an AI-powered web application that summarizes videos and audio from platforms like Bilibili, YouTube, TikTok, and podcasts, gene… | 57 | 6194 | active |
| svc-develop-team/so-vits-svc A deep learning framework based on SoftVC VITS for singing voice conversion (SVC), letting users train models that convert one singing voic… | 10 | 28125 | maintenance |
| bytedance/LatentSync LatentSync is an end-to-end lip-sync framework from ByteDance based on audio-conditioned latent diffusion models, using Stable Diffusion to… | 33 | 6026 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| modstart-lib/aigcpanel AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, … | 88 | 5482 | active |
| lemonade-sdk/lemonade Lemonade is a local AI server that runs optimized LLMs (plus image, speech, and embedding models) on your own GPU and NPU, exposing OpenAI-… | 82 | 5472 | active |
| signalwire/freeswitch FreeSWITCH is an open-source software-defined telecom stack that turns commodity servers into a full telephony platform for voice, video, a… | 94 | 5118 | stable |
| Plachtaa/VITS-fast-fine-tuning A Python pipeline for fast fine-tuning of VITS text-to-speech models, enabling speaker adaptation in under an hour from short audio, long a… | 10 | 5012 | active |
| PennyroyalTea/gibberlink GibberLink is a viral demo application in which two conversational AI agents (built on ElevenLabs) detect that they are both AI and switch … | 35 | 4887 | active |
| Osmantic/ODS ODS (Osmantic Deployment System) is a self-hosted AI server installer and runtime that turns a PC, Mac, or Linux machine into a private AI … | 80 | 4716 | active |
| OpenNMT/CTranslate2 CTranslate2 is a C++ and Python library for fast, memory-efficient inference of Transformer models on CPU and GPU. It uses quantization, la… | 96 | 4644 | stable |
| Rikorose/DeepFilterNet DeepFilterNet is a low-complexity speech enhancement framework that performs real-time noise suppression on full-band 48kHz audio using dee… | 23 | 4632 | active |
| crmne/ruby_llm RubyLLM is a Ruby framework providing a unified, expressive interface to all major AI providers (OpenAI, Anthropic, Google, Ollama, and any… | 83 | 4322 | active |
| Spark NLP Spark NLP is an open-source natural language processing library built natively on Apache Spark, providing scalable NLP annotations and tran… | 96 | 4159 | stable |
| QwenLM/Qwen2.5-Omni Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre… | 31 | 4074 | active |
| tensorflow/tensor2tensor Tensor2Tensor (T2T) is a Python library of deep learning models and datasets built on TensorFlow, developed by the Google Brain team to mak… | 10 | 17464 | maintenance |
| OpenMinis/OpenMinis OpenMinis is a free, open-source mobile AI agent app for iOS and Android that connects leading LLM providers (Claude, GPT, Gemini, etc.) vi… | 69 | 3985 | active |
| claraverse-space/ClaraVerse ClaraVerse is a self-hosted, privacy-focused AI workspace that combines chat, multi-agent teams, a visual workflow builder, RAG document pr… | 80 | 3894 | active |
| Plachtaa/seed-vc Seed-VC is a Python tool and model for zero-shot voice conversion, real-time voice conversion, and singing voice conversion, cloning a voic… | 10 | 3888 | active |
| PeterH0323/Streamer-Sales Streamer-Sales is a fine-tuned LLM-based AI sales livestreamer application that generates persuasive product commentary from product descri… | 27 | 3761 | active |
| HeartMuLa/heartlib HeartMuLa is a family of open-source music foundation models that generate music conditioned on lyrics and tags with multilingual support. … | 50 | 3749 | active |
| fudan-generative-vision/hallo2 Hallo2 is a Python research library from Fudan University that animates a single portrait image using audio input, producing long-duration … | 26 | 3734 | active |
| OvidijusParsiunas/deep-chat Deep Chat is a fully customizable, framework-agnostic AI chatbot web component that can be added to any website with one line of code. It s… | 91 | 3704 | active |
| sugarforever/chat-ollama ChatOllama is an open-source, self-hosted AI chatbot platform built with Nuxt 3 that supports multiple LLM providers (OpenAI, Anthropic, Ge… | 63 | 3509 | active |
| gnekt/My-Brain-Is-Full-Crew A crew of 8+ AI agents with 14 specialized skills that manages an Obsidian vault through conversational chat, handling knowledge organizati… | 55 | 3466 | active |
| huangjunsen0406/py-xiaozhi py-xiaozhi is an open-source, cross-platform multimodal AI voice assistant client written in Python, compatible with the xiaozhi-esp32 ecos… | 85 | 3452 | active |
| Grt1228/chatgpt-java An unofficial Java SDK for the OpenAI API covering all official endpoints including chat completions (GPT-3.5/GPT-4), DALL-E image generati… | 21 | 3423 | active |
| nicedreamzapp/claude-code-local An MLX-native server for Apple Silicon that exposes an Anthropic-API-compatible endpoint so Claude Code and similar agents can run entirely… | 77 | 3240 | active |
| alexrudall/ruby-openai A Ruby gem providing a client for the OpenAI API, supporting chat completions, the Responses API, streaming, vision, audio (Whisper), image… | 72 | 3222 | active |
| Rudrabha/Wav2Lip Wav2Lip is the official research code for the ACM Multimedia 2020 paper 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In … | 45 | 13182 | maintenance |
| Snailclimb/interview-guide An open-source AI-powered interview platform built with Spring Boot 4.1, Java 25, Spring AI 2.0, React, PostgreSQL/pgvector, and Redis. It … | 59 | 3106 | active |
| chaterm/Chaterm Chaterm is an open-source AI-native terminal application for cloud and infrastructure management, letting engineers deploy, troubleshoot, a… | 80 | 3029 | active |
| TanStack/ai TanStack AI is a type-safe, provider-agnostic TypeScript SDK for building AI applications with streaming chat, tool calling, agents, struct… | 80 | 3028 | active |
| SharpAI/DeepCamera DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r… | 86 | 3019 | active |
| zcaceres/markdownify-mcp A Model Context Protocol (MCP) server that converts PDFs, images, audio, Office documents, and web content (including YouTube transcripts a… | 81 | 2983 | active |
| MacPaw/OpenAI A community-maintained Swift package that wraps the OpenAI public API, supporting chat completions, responses, function calling, MCP tools,… | 94 | 2935 | active |
| JunChenMoCode/ChatGPT_JCM A Vue2 + ElementUI web management interface that aggregates OpenAI API endpoints (models, chat, images, audio, fine-tuning, files) into a g… | 22 | 2923 | active |
| datascale-ai/opentalking OpenTalking is an open-source Python framework for building real-time AI digital-human (talking avatar) conversation products. It orchestra… | 65 | 2897 | active |
| luoluoluo22/jianying-editor-skill A Python-based AI agent skill that automates video editing in JianYing (the Chinese desktop version of CapCut) by directly generating and i… | 55 | 2878 | active |
| sahibzada-allahyar/YC-Killer A collection of open-source, enterprise-grade AI agents intended as free alternatives to commercial Y Combinator startups. It includes a de… | 64 | 2792 | active |
| QwenLM/Qwen-MM-Plugins A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.… | 57 | 2777 | active |
| hoowhoami/EchoMusic EchoMusic is an open-source, minimalist third-party desktop music player client for Kugou Concept Edition, built with Electron, Vue 3, Type… | 82 | 2729 | active |
| prophesier/diff-svc Diff-SVC is a deep learning project that performs singing voice conversion using diffusion models, transforming input singing audio into a … | 62 | 2717 | active |
| Project-N-E-K-O/N.E.K.O Project N.E.K.O. is an open-source AI companion application — a proactive catgirl-style AI that lives on your desktop, initiates interactio… | 85 | 2681 | active |
| openai/openai-dotnet The official .NET library for accessing the OpenAI REST API, generated from the OpenAPI specification in collaboration with Microsoft. It p… | 86 | 2674 | active |
| heshengtao/super-agent-party Super Agent Party is a self-hosted, all-in-one AI desktop companion that combines VRM-based virtual characters, agent skills, MCP tool supp… | 87 | 2607 | active |
| asteroid-team/asteroid Asteroid is a PyTorch-based audio source separation toolkit for researchers, providing modular building blocks (filterbanks, encoders, mask… | 60 | 2584 | active |
| mozilla/TTS A deep learning library for advanced text-to-speech generation, built on PyTorch with models like Tacotron2, Glow-TTS, and various vocoders… | 23 | 10167 | maintenance |
| VITA-MLLM/VITA VITA is an open-source interactive omni multimodal large language model (VITA-1.5) that supports real-time vision and speech interaction, s… | 29 | 2534 | active |
| NeuralNomadsAI/CodeNomad CodeNomad is a premium desktop application that wraps the OpenCode AI coding agent into a full graphical workspace with multi-instance sess… | 83 | 2514 | active |
| aurora-develop/aurora A Go service that exposes ChatGPT Web capabilities as an OpenAI-compatible API, including chat completions, responses, file Q&A, image gene… | 90 | 2485 | active |
| volcengine/ai-app-lab AI App Lab from Volcano Engine (Volcengine) provides Arkitect, a high-code Python SDK for building LLM applications, plus Demohouse, a coll… | 83 | 2433 | active |
| Alibaba-Quark/LiveAvatar LiveAvatar is an open-source implementation of an ECCV 2026 paper for streaming, real-time, infinite-length audio-driven avatar video gener… | 61 | 2386 | active |
| ailia-ai/ailia-models A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,… | 77 | 2385 | active |
| vercel/ai-elements AI Elements is a React component library and custom shadcn/ui registry from Vercel providing pre-built, composable UI components for AI-nat… | 75 | 2365 | active |
| snapotter-hq/SnapOtter SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities… | 80 | 2362 | active |
| heshengtao/comfyui_LLM_party A ComfyUI plugin providing a comprehensive set of nodes for building LLM agent workflows, including MCP server support, RAG/GraphRAG, TTS, … | 65 | 2343 | active |
| stephengpope/no-code-architects-toolkit A self-hostable Flask-based API that consolidates common media processing tasks—video editing, captioning, audio conversion, transcription,… | 50 | 2339 | active |
| codedogQBY/ReadAny ReadAny is a local-first, AI-powered cross-platform e-book reader for desktop and mobile. It combines RAG-based chat with your books, hybri… | 81 | 2267 | active |
| rushindrasinha/youtube-shorts-pipeline Verticals v3 (repo youtube-shorts-pipeline) is a Python CLI that automates producing and publishing YouTube Shorts: it researches a topic, … | 61 | 2251 | active |
| Alpha-VLLM/Lumina-T2X Lumina-T2X is a unified framework for text-to-any-modality generation built on flow-based large diffusion transformers. It supports generat… | 28 | 2250 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2203 | active |
| TNT-Likely/BeeCount BeeCount is an open-source, local-first personal bookkeeping app for iOS, Android, and Web built with Flutter, offering multi-ledger, multi… | 80 | 2188 | active |
| walterlow/freecut FreeCut is a professional-grade, browser-based multi-track video editor requiring no installation or uploads, with all media and project fi… | 60 | 2096 | active |
| vitoplantamura/OnnxStream A lightweight C++ inference library for ONNX models that streams weights to run large models in very little memory, accelerated by XNNPACK.… | 59 | 2086 | active |
| MiniMax-AI/cli The official CLI for the MiniMax AI Platform, written in TypeScript, that generates text, images, video, speech, and music from the termina… | 81 | 2073 | active |
| Kochava-Studios/witsy Witsy is a cross-platform desktop AI assistant built with Electron and Vue 3 that lets users chat with many LLM providers using their own A… | 69 | 2024 | active |
| 64bit/async-openai async-openai is a Rust library providing typed, async clients for the OpenAI API, covering chat completions, responses, embeddings, assista… | 93 | 1998 | active |
| crow-translate/crow-translate Crow Translate is a lightweight C++/Qt desktop translator that translates and speaks text using Google, Yandex, Bing, LibreTranslate, and L… | 10 | 1978 | active |
| run-llama/notebookllama NotebookLlaMa is an open-source, Python-based alternative to Google's NotebookLM, backed by LlamaCloud for document ingestion and retrieval… | 52 | 1967 | active |
| StanfordBDHG/HealthGPT HealthGPT is an experimental open-source iOS app built on the Stanford Spezi framework that lets users query their Apple Health data using … | 83 | 1964 | active |
| alexpinel/Dot Dot is a standalone Electron desktop application for chatting with your documents using fully local LLMs and Retrieval Augmented Generation… | 16 | 1911 | active |
| FACEGOOD/FACEGOOD-Audio2Face FACEGOOD Audio2Face is an open-source deep learning framework that converts audio into facial blendshape weights for driving digital humans… | 64 | 1909 | active |
| szczyglis-dev/py-gpt PyGPT is an open-source, all-in-one desktop AI assistant for Linux, Windows, and Mac, written in Python. It supports chat, agents, vision, … | 92 | 1892 | active |