function: audio-processing
1675 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| yt-dlp/yt-dlp yt-dlp is a feature-rich command-line audio/video downloader supporting thousands of websites, forked from youtube-dl. It offers extensive … | 99 | 187167 | active |
| ComfyUI ComfyUI is a modular, node-graph based GUI, API, and backend for running diffusion models and other generative AI models to create images, … | 94 | 130155 | stable |
| OpenCut-app/OpenCut OpenCut is a free, open-source video editor for web, desktop, and mobile, positioned as an open alternative to CapCut. It provides a timeli… | 78 | 86959 | active |
| RVC-Boss/GPT-SoVITS GPT-SoVITS is a Python-based few-shot voice cloning and text-to-speech system with an integrated WebUI. It supports zero-shot TTS from a 5-… | 68 | 61255 | active |
| jamiepine/voicebox Voicebox is a free, open-source, local-first AI voice studio desktop app that combines voice cloning, text-to-speech across 7 engines, and … | 75 | 51551 | active |
| 9001/copyparty A portable, single-file Python file server that turns almost any device into a NAS-like file sharing server with accelerated resumable uplo… | 95 | 46408 | active |
| LizardByte/Sunshine Sunshine is a self-hosted game stream host that works with Moonlight clients, providing low-latency cloud gaming server capabilities. It su… | 90 | 40571 | active |
| myshell-ai/OpenVoice OpenVoice is a Python library and audio foundation model for instant voice cloning, requiring only a short reference audio clip to replicat… | 34 | 37312 | active |
| ZuodaoTech/everyone-can-use-english Everyone Can Use English is an open-source project combining a well-known English learning book with the Enjoy app, an AI-powered language … | 74 | 36774 | active |
| google-ai-edge/mediapipe MediaPipe is Google's cross-platform framework for deploying on-device machine learning solutions for live and streaming media. It provides… | 94 | 36731 | stable |
| OpenBMB/VoxCPM VoxCPM is a tokenizer-free text-to-speech system built on a diffusion autoregressive architecture that generates continuous speech represen… | 79 | 36145 | active |
| raysan5/raylib raylib is a simple and easy-to-use C library for videogames programming, providing bindings over OpenGL for 2D/3D graphics, input, audio, a… | 81 | 34476 | stable |
| huggingface/diffusers Hugging Face Diffusers is a Python library providing state-of-the-art pretrained diffusion models for generating images, videos, and audio … | 95 | 34385 | stable |
| qarmin/czkawka Czkawka is a fast, multiplatform app written in Rust that finds duplicate files, empty folders, similar images, similar videos, similar mus… | 88 | 33012 | active |
| telegramdesktop/tdesktop The official open-source Telegram Desktop messenger client, built on the Telegram API and MTProto secure protocol. It is a full-featured cr… | 95 | 32746 | stable |
| fishaudio/fish-speech Fish Speech is an open-source state-of-the-art text-to-speech system (Fish Audio S2) trained on over 10 million hours of audio across ~50 l… | 74 | 32413 | active |
| Zackriya-Solutions/meetily Meetily is a privacy-first, open-source AI meeting assistant that records, transcribes, and summarizes meetings entirely locally using Para… | 79 | 29940 | active |
| Signal Private Messenger Signal is a free, open-source private messenger offering end-to-end encrypted text, voice, and video communication, built on the Signal Pro… | 95 | 29262 | stable |
| 78/xiaozhi-esp32 XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen… | 88 | 29186 | active |
| ossrs/srs SRS (Simple Realtime Server) is a high-performance, open-source real-time media server written in C++ that supports RTMP, WebRTC, HLS, HTTP… | 98 | 29169 | stable |
| Label Studio Label Studio is an open-source data labeling and annotation platform supporting images, audio, text, video, and time series with a web UI a… | 86 | 28150 | active |
| BabylonJS/Babylon.js Babylon.js is an open-source 3D game and rendering engine for the web, written in TypeScript and built on WebGL, WebGPU, and Web Audio. It … | 95 | 25985 | stable |
| copy/v86 v86 is an x86 PC emulator written in JavaScript that translates x86 machine code to WebAssembly at runtime via a JIT compiler, running enti… | 86 | 23424 | active |
| QwenAudio/CosyVoice CosyVoice is a multilingual large voice generation model (TTS) with full-stack inference, training, and deployment support. It offers zero-… | 61 | 22925 | active |
| huggingface/datasets Hugging Face Datasets is a Python library providing one-line access to hundreds of thousands of public datasets on the Hugging Face Hub acr… | 98 | 21870 | stable |
| bluenviron/mediamtx MediaMTX is a zero-dependency live media server and proxy written in Go that routes real-time video and audio streams across protocols incl… | 99 | 19938 | stable |
| huggingface/sentence-transformers Sentence Transformers (SBERT) is a Python library for computing, using, and training state-of-the-art embedding, reranker, sparse encoder, … | 99 | 19036 | stable |
| jianchang512/pyvideotrans pyVideoTrans is an open-source desktop application (with WebUI and CLI modes) that translates videos from one language to another. It provi… | 90 | 18812 | active |
| NVIDIA-NeMo/Speech NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, … | 98 | 18337 | active |
| Pion WebRTC Pion WebRTC is a pure Go implementation of the WebRTC API, enabling real-time audio, video, and data communication without cgo or external … | 98 | 16742 | active |
| huggingface/transformers.js Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) … | 91 | 16270 | active |
| xournalpp/xournalpp Xournal++ is a cross-platform handwriting notetaking application written in C++ with GTK3, supporting pressure-sensitive stylus input from … | 94 | 15286 | active |
| SWivid/F5-TTS F5-TTS is the official implementation of a fully non-autoregressive text-to-speech system based on flow matching with a Diffusion Transform… | 85 | 15167 | active |
| kekingcn/kkFileView kkFileView is a self-hosted Spring Boot application that provides online preview of a very wide range of file formats, including Office doc… | 95 | 14595 | active |
| Audiobookshelf Audiobookshelf is an open-source, self-hosted media server for streaming and managing audiobooks and podcasts, with companion Android and i… | 97 | 14135 | active |
| AlexxIT/go2rtc go2rtc is a zero-dependency camera streaming application and media server written in Go that supports dozens of streaming formats and proto… | 84 | 14050 | active |
| SubtitleEdit/subtitleedit Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac… | 94 | 13971 | active |
| Omi Omi is an open-source AI wearable and companion app ecosystem (necklace pendant, smart glasses, desktop and mobile apps) that captures conv… | 88 | 13264 | active |
| DustinBrett/daedalOS daedalOS is a desktop environment that runs entirely in the browser, built with JavaScript. It provides a window manager, file system, task… | 77 | 13025 | active |
| modelscope/DiffSynth-Studio DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio … | 79 | 13003 | active |
| abus-aikorea/voice-pro Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, … | 84 | 12647 | active |
| ace-step/ACE-Step-1.5 ACE-Step 1.5 is an open-source music generation foundation model combining a language model planner with a Diffusion Transformer to create … | 79 | 12421 | active |
| cleanlab/cleanlab Cleanlab is a Python library for data-centric AI that automatically detects issues in ML datasets, such as label errors, outliers, duplicat… | 62 | 11636 | stable |
| kyutai-labs/moshi Moshi is a speech-text foundation model and full-duplex spoken dialogue framework from Kyutai, built around the Mimi streaming neural audio… | 62 | 10949 | active |
| Baiyuetribe/paper2gui Paper2GUI (now branded 小白兔AI / Xiaobaitu AI) is a desktop AI toolbox application that packages 50+ AI research models into install-free GUI… | 23 | 10709 | active |
| nexmoe/VidBee VidBee is a free, open-source Electron desktop app for downloading video and audio from 1000+ sites (via yt-dlp) and importing local media.… | 84 | 10386 | active |
| espnet/espnet ESPnet is an end-to-end speech processing toolkit built on PyTorch covering speech recognition, text-to-speech, speech translation, enhance… | 88 | 9941 | active |
| FyroxEngine/Fyrox Fyrox is a feature-rich, general-purpose 2D and 3D game engine written in Rust, formerly known as rg3d. It includes a scene editor, GUI sys… | 83 | 9525 | active |
| k2-fsa/OmniVoice OmniVoice is a massively multilingual zero-shot text-to-speech model supporting 600+ languages, built on a diffusion language model-style a… | 79 | 9455 | active |
| chrismaltby/gb-studio GB Studio is a drag-and-drop visual game creator for making Game Boy games without programming knowledge, built as an Electron desktop app … | 96 | 9385 | active |
| LTX-2 Official Python package from Lightricks providing inference pipelines and LoRA training for LTX-2/LTX-2.5, an open-weights DiT-based founda… | 83 | 9260 | active |
| lipku/LiveTalking LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit… | 89 | 9238 | active |
| bigbluebutton/bigbluebutton BigBlueButton is an open-source web conferencing system purpose-built for online learning, supporting real-time audio, video, slides with w… | 98 | 9202 | active |
| coqui-ai/TTS Coqui TTS is a deep learning toolkit for text-to-speech synthesis, providing pretrained models in over 1100 languages plus tools for traini… | 23 | 45953 | maintenance |
| fastrepl/anarlog Anarlog is an open-source, local-first AI meeting notepad and Granola alternative that records device audio without a bot joining the call,… | 84 | 9181 | active |
| meetecho/janus-gateway Janus is an open-source, general-purpose WebRTC server written in C that sets up media communication with browsers and relays RTP/RTCP and … | 75 | 9156 | stable |
| fudan-generative-vision/hallo Hallo is a Python research library implementing hierarchical audio-driven visual synthesis for animating portrait images into talking-head … | 14 | 8664 | active |
| nl8590687/ASRT_SpeechRecognition ASRT is a deep-learning-based Chinese speech recognition (speech-to-text) system built with TensorFlow/Keras, using CNN, LSTM, attention me… | 57 | 8383 | active |
| duixcom/Duix-Mobile Duix Mobile is an open-source SDK for building real-time interactive AI avatars (digital humans) that run on-device on Android, iOS, tablet… | 69 | 8198 | active |
| jianchang512/ChatTTS-ui A local web interface for the ChatTTS text-to-speech model that synthesizes speech from mixed Chinese/English text, numbers, and symbols. I… | 70 | 7640 | active |
| babysor/MockingBird MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real… | 54 | 36909 | maintenance |
| mediasoup mediasoup is a powerful WebRTC SFU (Selective Forwarding Unit) implemented as a Node.js module and Rust crate, with a C++ media worker subp… | 95 | 7345 | stable |
| ilyhalight/voice-over-translation A userscript/browser extension that adds AI voice-over translation and subtitles to videos on many websites, powered by Yandex services. It… | 98 | 7335 | active |
| turanszkij/WickedEngine Wicked Engine is an open-source C++ 3D game engine with modern graphics features like ray tracing, global illumination, and physically base… | 99 | 7203 | active |
| gaozhangmin/boxplayer BoxPlayer is a cross-platform desktop application that unifies multiple cloud drives (Aliyun Drive, Baidu, 115, Quark, OneDrive, etc.), loc… | 97 | 6878 | active |
| espeak-ng/espeak-ng eSpeak NG is a compact open-source text-to-speech synthesizer supporting over 100 languages and accents, using formant synthesis for small … | 67 | 6763 | active |
| ml5js/ml5-library ml5.js is a friendly, beginner-oriented JavaScript machine learning library for the browser, built on top of TensorFlow.js. It provides acc… | 23 | 6587 | active |
| makepad/makepad Makepad is a creative software development platform for Rust that provides a cross-platform UI runtime with a live-editable design DSL, com… | 67 | 6580 | active |
| netless-io/flat Flat is the open-source Web, Windows, and macOS client of Agora Flat, an online classroom platform. It provides real-time interactive white… | 70 | 6415 | active |
| vllm-project/vllm-omni vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer… | 83 | 6369 | active |
| nesbox/TIC-80 TIC-80 is a free, open-source fantasy computer for making, playing, and sharing tiny retro games. It bundles code, sprite, map, and sound e… | 64 | 6103 | active |
| jindrapetrik/jpexs-decompiler JPEXS Free Flash Decompiler (FFDec) is an open-source Java application for decompiling and editing Adobe Flash SWF files. It extracts resou… | 93 | 5828 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| CharlesPikachu/musicdl A lightweight music downloader written in pure Python that supports dozens of music and audiobook platforms including NetEase Cloud Music, … | 75 | 5757 | active |
| Eventual-Inc/Daft Daft is a high-performance distributed data engine with a Python dataframe API, implemented in Rust, designed for AI and multimodal workloa… | 95 | 5730 | active |
| sohzm/cheating-daddy Cheating Daddy is a free, open-source Electron desktop app that acts as a real-time AI assistant during video calls, interviews, and meetin… | 78 | 5580 | active |
| modstart-lib/aigcpanel AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, … | 88 | 5482 | active |
| mifi/editly Editly is a declarative non-linear video editing tool and framework built on Node.js and ffmpeg, offering both a CLI and a JavaScript API. … | 32 | 5477 | active |
| remsky/Kokoro-FastAPI A Dockerized FastAPI wrapper around the Kokoro-82M text-to-speech model exposing an OpenAI-compatible speech endpoint with CPU, NVIDIA, AMD… | 88 | 5373 | active |
| liuzhao1225/YouDub-webui YouDub WebUI is an open-source AI video localization and dubbing tool that converts YouTube, Bilibili, or local videos into target-language… | 75 | 5358 | active |
| mosra/magnum Magnum is a lightweight and modular C++11/C++14 graphics middleware library providing a slim wrapper over OpenGL, OpenGL ES, WebGL and Vulk… | 67 | 5194 | stable |
| signalwire/freeswitch FreeSWITCH is an open-source software-defined telecom stack that turns commodity servers into a full telephony platform for voice, video, a… | 94 | 5118 | stable |
| mediacms-io/mediacms MediaCMS is an open source, self-hosted video and media CMS built with Django and React, exposing a REST API. It supports video, audio, ima… | 99 | 5083 | active |
| tixl3d/tixl TiXL (formerly Tooll3) is a free, open-source desktop application for creating realtime motion graphics, combining graph-based procedural c… | 90 | 5081 | active |
| RandyGaul/cute_headers A collection of cross-platform, single-file C/C++ header-only libraries with no dependencies, primarily aimed at game development. It inclu… | 75 | 5052 | active |
| Plachtaa/VITS-fast-fine-tuning A Python pipeline for fast fine-tuning of VITS text-to-speech models, enabling speaker adaptation in under an hour from short audio, long a… | 10 | 5012 | active |
| Emby Server Emby Server is a self-hosted personal media server that organizes videos, music, photos, and live TV into a rich library and streams them t… | 23 | 4961 | active |
| microsoft/muzic Muzic is a Microsoft Research project providing deep learning models for music understanding and generation, including MusicBERT, SongMASS,… | 66 | 4952 | active |
| neilsonnn/image-blaster A Claude skillset that converts a single input image into a full 3D environment, including meshed 3D models (.glb/.obj), Gaussian splats (.… | 51 | 4822 | active |
| snesrev/zelda3 A reverse-engineered reimplementation of The Legend of Zelda: A Link to the Past written in ~70-80k lines of C, playable from start to fini… | 23 | 4739 | active |
| ant-media/Ant-Media-Server Ant Media Server is a scalable open-source media server for ultra-low latency live streaming, built around WebRTC (~0.5s latency) with supp… | 89 | 4726 | stable |
| Steam-Headless/docker-steam-headless A headless Steam Docker image that runs a full Steam client and Xfce4 desktop on Linux with NVIDIA, AMD, or Intel GPU support. It enables r… | 58 | 4726 | active |
| miroslavpejic85/mirotalk MiroTalk P2P is a self-hosted, open-source WebRTC video conferencing platform that uses direct peer-to-peer connections for real-time video… | 77 | 4706 | active |
| KaijuEngine/kaiju Kaiju is an open-source 2D/3D game engine written in Go and backed by Vulkan, with a built-in visual editor that itself runs as a game in t… | 80 | 4694 | active |
| WhisperSpeech/WhisperSpeech WhisperSpeech is an open-source text-to-speech system built by inverting OpenAI's Whisper model, aiming to be 'Stable Diffusion for speech'… | 56 | 4639 | active |
| LaoFeng-mouse/flyingmouse-format FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop… | 79 | 4601 | active |
| Cocos2d-x Cocos2d-x is an open-source, cross-platform C++ framework for building 2D games and graphical applications, with Lua and JavaScript binding… | 41 | 19162 | maintenance |
| KilledByAPixel/LittleJS LittleJS is a tiny, fast, open-source HTML5 game engine written in JavaScript with no dependencies. It bundles WebGL2/Canvas2D rendering, a… | 94 | 4166 | active |
| QwenLM/Qwen2.5-Omni Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre… | 31 | 4074 | active |
| jitwxs/163MusicLyrics A Windows desktop GUI application for fetching and processing song lyrics from NetEase Cloud Music and QQ Music. It supports searching by s… | 74 | 4071 | active |