domain: speech-processing
552 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| modelscope/modelscope ModelScope is a Python library and ecosystem built on the 'Model-as-a-Service' concept, providing unified APIs to download, run inference o… | 98 | 9111 | active |
| duixcom/Duix-Mobile Duix Mobile is an open-source SDK for building real-time interactive AI avatars (digital humans) that run on-device on Android, iOS, tablet… | 69 | 8198 | active |
| fonoster/fonoster Fonoster is an open-source programmable telecommunications stack (a Twilio alternative) for building voice and messaging applications, with… | 99 | 8080 | active |
| facebookresearch/fairseq Fairseq is a PyTorch-based sequence modeling toolkit from Facebook AI Research for training custom models for translation, summarization, l… | 10 | 32231 | maintenance |
| vllm-project/vllm-omni vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer… | 83 | 6369 | active |
| Shaunwei/RealChar RealChar is an open-source application for creating, customizing, and talking to AI characters/companions in realtime via voice or text. It… | 48 | 6213 | active |
| cactus-compute/cactus Cactus is a hybrid edge-cloud AI inference engine for mobile devices, wearables, smart home devices, and robots, built in C++ with custom q… | 86 | 5934 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| pytorch/executorch ExecuTorch is PyTorch's framework for exporting and running AI models on-device across mobile, embedded, and edge hardware, with a tiny (~5… | 95 | 4953 | active |
| OpenNMT/CTranslate2 CTranslate2 is a C++ and Python library for fast, memory-efficient inference of Transformer models on CPU and GPU. It uses quantization, la… | 96 | 4644 | stable |
| tensorflow/tensor2tensor Tensor2Tensor (T2T) is a Python library of deep learning models and datasets built on TensorFlow, developed by the Google Brain team to mak… | 10 | 17464 | maintenance |
| PeterH0323/Streamer-Sales Streamer-Sales is a fine-tuned LLM-based AI sales livestreamer application that generates persuasive product commentary from product descri… | 27 | 3761 | active |
| HumanAIGC-Engineering/OpenAvatarChat OpenAvatarChat is a modular interactive digital human (talking avatar) chat application that combines ASR, LLM, TTS, and avatar rendering c… | 74 | 3723 | active |
| sonos/tract Tract is Sonos' tiny, self-contained neural-network inference engine written in Rust. It loads ONNX, TensorFlow/TFLite, and NNEF models, op… | 99 | 3045 | active |
| datascale-ai/opentalking OpenTalking is an open-source Python framework for building real-time AI digital-human (talking avatar) conversation products. It orchestra… | 65 | 2897 | active |
| sahibzada-allahyar/YC-Killer A collection of open-source, enterprise-grade AI agents intended as free alternatives to commercial Y Combinator startups. It includes a de… | 64 | 2792 | active |
| iamsrikanthnani/pluely Pluely is a privacy-first desktop AI assistant that runs as an invisible always-on-top overlay, providing live meeting transcription, on-sc… | 79 | 2596 | active |
| VITA-MLLM/VITA VITA is an open-source interactive omni multimodal large language model (VITA-1.5) that supports real-time vision and speech interaction, s… | 29 | 2534 | active |
| Intent-Lab/VisionClaw VisionClaw is a real-time AI assistant app for Meta Ray-Ban smart glasses that streams camera frames and microphone audio to the Gemini Liv… | 59 | 2529 | active |
| apple/axlearn AXLearn is a Python deep learning library built on JAX and XLA for developing and training large-scale models, with an object-oriented conf… | 71 | 2372 | active |
| Natively-AI-assistant/natively-cluely-ai-assistant Natively is a free, source-available desktop AI meeting assistant and interview copilot that provides real-time transcription, AI-generated… | 82 | 2354 | active |
| espressif/esp-adf Espressif's official Advanced Development Framework for building audio and multimedia applications on ESP32-series SoCs. It provides pipeli… | 80 | 2299 | active |
| floneum/kalosm Kalosm is a Rust ecosystem of crates providing simple interfaces for running pre-trained language, audio, and image models locally or remot… | 64 | 2223 | active |
| vitoplantamura/OnnxStream A lightweight C++ inference library for ONNX models that streams weights to run large models in very little memory, accelerated by XNNPACK.… | 59 | 2086 | active |
| ROCm/FastFlowLM FastFlowLM (FLM) is an NPU-first LLM inference runtime purpose-built and deeply optimized for AMD Ryzen AI NPUs (XDNA2), offering an Ollama… | 85 | 1809 | active |
| tuya/TuyaOpen TuyaOpen is an open-source, cross-platform C/C++ SDK and IoT OS for building AI-agent hardware on Tuya T-series MCUs, ESP32, Beken, LN882H,… | 87 | 1804 | active |
| software-mansion/react-native-executorch React Native ExecuTorch is a declarative React Native library for running AI models on-device, powered by Meta's ExecuTorch runtime. It shi… | 88 | 1702 | active |
| semperai/amica Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. … | 32 | 1593 | active |
| TOM88812/xiaozhi-android-client A Flutter-based cross-platform voice chat client for the xiaozhi AI assistant ecosystem, supporting real-time voice interaction and text co… | 60 | 1574 | active |
| waybarrios/vllm-mlx vllm-mlx is a vLLM-style LLM inference server for Apple Silicon Macs built on native MLX, exposing both OpenAI and Anthropic compatible API… | 82 | 1547 | active |
| dindin0497/SeeIt SeeIt is an inclusive Android app with two accessibility modes: one that converts spoken speech into text and plays corresponding ASL (Amer… | 41 | 1540 | active |
| sauravpanda/BrowserAI BrowserAI is a TypeScript library for running LLMs, speech recognition, text-to-speech, and audio separation models directly in the browser… | 78 | 1449 | active |
| ARahim3/mlx-tune A Python library for fine-tuning LLMs, vision-language, audio (TTS/STT), embedding, OCR, and JEPA models natively on Apple Silicon Macs usi… | 75 | 1389 | active |
| elevenyellow/handcrafted-persona-engine Persona Engine is a Windows desktop application that drives a Live2D avatar with an AI pipeline: microphone speech recognition, an LLM guid… | 73 | 1357 | active |
| morettt/my-neuro An open-source AI desktop companion framework inspired by Neuro-sama, letting users build a customizable Live2D character with sub-second v… | 84 | 1342 | active |
| Capsize-Games/airunner AI Runner is a privacy-focused desktop application for running local AI models offline, combining an AI chat companion with voice conversat… | 79 | 1314 | active |
| espressif/esp-box ESP-BOX is Espressif's AIoT development framework for the ESP32-S3-BOX series of development boards, built on the ESP32-S3 Wi-Fi + Bluetoot… | 61 | 1297 | active |
| kubeai-project/kubeai KubeAI is a Kubernetes operator for serving machine learning models in production, supporting LLMs via vLLM and Ollama, vector embeddings, … | 89 | 1256 | active |
| GML-MMGroup/GMTalker GMTalker is an interactive 3D digital human system rendered with Unreal Engine, integrating speech recognition, speech synthesis, natural l… | 46 | 1217 | active |
| uezo/ChatdollKit ChatdollKit is a Unity SDK that turns 3D character models into voice-enabled chatbots and virtual assistants. It integrates LLMs (ChatGPT, … | 76 | 1213 | active |
| HITsz-TMG/Uni-MoE Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,… | 68 | 1116 | active |
| nobodywho-ooo/nobodywho NobodyWho is an open-source (EUPL-1.2) on-device LLM inference engine written in Rust, built on llama.cpp, with SDKs for Kotlin, Swift, Pyt… | 86 | 1086 | active |
| susiai/susi_chat A chat interface for communicating with a locally hosted LLM (via llama.cpp) through a terminal console, a browser-based console, or a voic… | 61 | 1036 | active |
| microsoft/torchscale A PyTorch library from Microsoft implementing foundation Transformer architectures such as DeepNet, Magneto, RetNet, LongNet, BitNet, and X… | 32 | 3138 | maintenance |
| microsoft/i-Code Microsoft's i-Code is a collection of research models and frameworks for integrative, composable multimodal AI spanning vision, language, a… | 32 | 1703 | maintenance |
| yakGPT/yakGPT YakGPT is a locally running, browser-based ChatGPT UI that connects directly to the OpenAI API with your own key. It adds hands-free voice … | 30 | 1586 | maintenance |
| ARM-software/ML-KWS-for-MCU TensorFlow models and training scripts for keyword spotting (wake-word detection) on Arm Cortex-M microcontrollers, accompanying the 'Hello… | 32 | 1249 | maintenance |
| Syan-Lin/CyberWaifu CyberWaifu is a Python chatbot application that combines LLMs (ChatGPT, Claude) with TTS (edge-tts, Azure) to create realistic conversation… | 30 | 1099 | maintenance |
| susiai/susi_device SUSI Device provides sources to install the SUSI AI assistant stack on a Raspberry Pi, combining microphone/speaker, a small display, a loc… | 32 | 1004 | maintenance |
| evilsocket/cake Cake is a multimodal AI inference server written in Rust that runs text, image, and voice models on a single device or shards them across a… | 59 | 3114 | experimental |
| C-Nedelcu/talk-to-chatgpt A Chrome and Edge browser extension that lets users talk to ChatGPT using speech recognition and hear responses via text-to-speech, with op… | 32 | 1929 | abandoned |
| yodaos-project/yodaos YodaOS is a Linux distribution built on OpenWrt for voice-enabled IoT devices, using JavaScript as its primary application language. It tar… | 32 | 1225 | abandoned |
← prev page 6 / 6