function: audio-processing
1675 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| SunoAI-API/Suno-API An unofficial Python/FastAPI-based API wrapper for Suno AI that lets you generate songs and lyrics programmatically. It includes automatic … | 33 | 1806 | active |
| FL33TW00D/whisper-turbo Whisper Turbo is a fast, cross-platform, GPU-accelerated implementation of OpenAI's Whisper speech recognition model, built on the Ratchet … | 19 | 1794 | active |
| HFrost0/bilix bilix is a lightning-fast asynchronous download tool for bilibili and other sites, usable via a CLI or as an extensible Python library. It … | 23 | 1787 | active |
| wwbin2017/bailing Bailing is an open-source voice assistant application similar to GPT-4o, built with an ASR + VAD + LLM + TTS pipeline (FunASR, silero-vad, … | 54 | 1757 | active |
| microsoft/DirectXTK12 DirectX Tool Kit for DirectX 12 is a C++ library of helper classes for writing Direct3D 12 code, covering sprite rendering, effects, textur… | 85 | 1749 | stable |
| elder-plinius/ST3GG ST3GG is an all-in-one steganography toolkit that hides secret data inside images, audio, documents, and network packets using 100+ encodin… | 63 | 1741 | active |
| NVIDIA-NeMo/Curator NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for … | 86 | 1736 | active |
| AutoArk/EVA-OS EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining … | 68 | 1712 | active |
| fdivitto/FabGL FabGL is a C++ graphics and input library for the ESP32, providing display drivers for VGA, PAL/NTSC composite video, and I2C/SPI LCDs, plu… | 26 | 1703 | active |
| OpenMOSS/MOSS-Transcribe-Diarize MOSS-Transcribe-Diarize 0.9B is an open-source end-to-end audio understanding model that jointly performs multi-speaker speech transcriptio… | 58 | 1695 | active |
| oakmound/oak Oak is a pure Go 2D game engine providing rendering, particles, collision, physics, audio, input, and event systems without external depend… | 66 | 1670 | active |
| TorqueGameEngines/Torque2D Torque2D is a free, MIT-licensed, open-source 2D game engine in C++ built on the proven GarageGames Torque technology. It offers cross-plat… | 90 | 1667 | active |
| Fantasy-AMAP/fantasy-talking FantasyTalking is a research codebase and model for generating realistic talking portrait videos from a single image and an audio clip, bui… | 48 | 1628 | active |
| ZiqiaoPeng/SyncTalk SyncTalk is the official PyTorch implementation of a CVPR 2024 paper that synthesizes speech-driven, synchronized talking head videos using… | 46 | 1626 | active |
| JohnEarnest/Decker Decker is a multimedia platform and sketchpad for creating and sharing interactive documents with sound, images, hypertext, and scripted be… | 76 | 1615 | active |
| hiloteam/Hilo Hilo is a cross-end HTML5 game development framework from Alibaba Group, providing a lightweight visual object architecture with multiple r… | 44 | 5939 | maintenance |
| ml4a/ml4a ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li… | 32 | 1602 | active |
| X-LANCE/AniTalker AniTalker is the official PyTorch implementation of an ACM MM 2024 paper that animates a single static portrait into a vivid talking-face v… | 24 | 1598 | active |
| akudamatata/Solara Solara is a minimalist web-based music player that aggregates multiple free music APIs for search, streaming playback, and audio downloads.… | 57 | 1582 | active |
| songloft-org/songloft Songloft (formerly MiMusic) is a free, ad-free, plugin-based self-hosted music server written in Go, designed for managing music you legall… | 82 | 1562 | active |
| jpush/aurora-imui Aurora IMUI is a general-purpose instant messaging UI component library providing MessageList and InputView components, independent of any … | 23 | 5696 | maintenance |
| shrimbly/node-banana Node Banana is an open-source, node-based visual workflow editor for building AI media generation pipelines. Users connect nodes on an infi… | 81 | 1550 | active |
| sepfy/libpeer libpeer is a portable WebRTC implementation written in C using BSD sockets, designed for IoT and embedded devices such as ESP32 and Raspber… | 77 | 1542 | active |
| vibevoice-community/VibeVoice VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker … | 60 | 1542 | active |
| tin2tin/Pallaidium Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE… | 75 | 1520 | active |
| Flashlight wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into … | 62 | 5466 | maintenance |
| Taiizor/Sucrose Sucrose is a free, open-source wallpaper engine for Windows that renders interactive live wallpapers from GIFs, videos, URLs, web pages, Yo… | 92 | 1503 | active |
| torch2424/wasmboy WasmBoy is a Game Boy and Game Boy Color emulator library written in WebAssembly using AssemblyScript, distributed via npm and importable i… | 32 | 1498 | active |
| QwenAudio/Fun-ASR Fun-ASR is a family of open-source LLM-based end-to-end speech recognition models from Tongyi Lab, covering Chinese, dialects, accents, and… | 81 | 1496 | active |
| p2r3/beheader A command-line tool that generates polyglot files - single files that are simultaneously valid images, videos, PDFs, ZIP archives, and HTML… | 48 | 1455 | active |
| silverstein/minutes Minutes is an open-source, privacy-first conversation memory app that records meetings, voice memos, and dictation, transcribes them locall… | 77 | 1453 | active |
| shridarpatil/whatomate Whatomate is an open-source, self-hosted WhatsApp Business platform shipped as a single Go binary with Docker support. It provides multi-te… | 69 | 1450 | active |
| voice-cloning-app/Voice-Cloning-App A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom… | 23 | 1440 | active |
| devnen/Chatterbox-TTS-Server A self-hosted server wrapping Resemble AI's Chatterbox TTS models behind an OpenAI-compatible API with a modern web UI. It supports voice c… | 65 | 1420 | active |
| K0lb3/UnityPy UnityPy is a Python library for extracting, unpacking, and editing Unity game engine asset files, based on AssetStudio. It supports reading… | 75 | 1414 | active |
| mpiorowski/late-sh late.sh is a terminal-first social platform accessed entirely over SSH, offering real-time chat, lofi/classical radio streaming, terminal g… | 77 | 1408 | active |
| Arthi-chaud/Meelo Meelo is a self-hosted music streaming server designed for music collectors, similar to Plex or Jellyfin but focused on music. It offers ri… | 99 | 1398 | active |
| gerbera/gerbera Gerbera is a free, open-source UPnP/DLNA media server that streams digital media across a home network to compatible devices like TVs, game… | 86 | 1390 | active |
| lamm-mit/PDF2Audio A Gradio-based web application that converts PDF documents into audio podcasts, lectures, and summaries using OpenAI GPT models for text ge… | 30 | 1383 | active |
| brainboxdotcc/DPP D++ (DPP) is a lightweight, modern C++ library for the Discord API, supporting API v10 with efficient caching, sharding, clustering, slash … | 98 | 1379 | stable |
| zelon88/HRConvert2 HRConvert2 is a self-hosted, resource-aware file conversion server written in PHP that supports 488 file formats across documents, images, … | 99 | 1364 | active |
| elevenyellow/handcrafted-persona-engine Persona Engine is a Windows desktop application that drives a Live2D avatar with an AI pipeline: microphone speech recognition, an LLM guid… | 73 | 1357 | active |
| yuanyuanxiang/SimpleRemoter SimpleRemoter (YAMA) is a C++ remote control suite derived from the Gh0st RAT codebase, providing remote desktop, file transfer, terminal, … | 60 | 1347 | active |
| AI-FanGe/OpenAIglasses_for_Navigation An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st… | 39 | 1343 | active |
| numandev1/react-native-compressor A React Native library that compresses images, videos, and audio with WhatsApp-like quality, plus background upload, file download, and vid… | 96 | 1325 | active |
| copperspice/copperspice CopperSpice is a set of cross-platform C++ libraries (Core, Gui, Network, Multimedia, SQL, OpenGL, Vulkan, WebKit, XML, and more) derived f… | 74 | 1324 | active |
| caiiiycuk/js-dos js-dos is a TypeScript library and API for running DOS and Windows 9x programs in the browser or Node.js, acting as a frontend for DOSBox/D… | 97 | 1322 | active |
| STMicroelectronics/STM32CubeF4 STMicroelectronics' official STM32Cube firmware package for the STM32F4 series of microcontrollers. It bundles HAL and LL drivers, CMSIS co… | 72 | 1316 | stable |
| DAMO-NLP-SG/VideoLLaMA2 VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro… | 25 | 1307 | active |
| StarlightSearch/EmbedAnything EmbedAnything is a high-performance, memory-safe embedding pipeline written in Rust (with Python bindings) that generates embeddings from t… | 88 | 1305 | active |
| Henry-23/VideoChat A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late… | 48 | 1303 | active |
| stenolabs/stenoai Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire… | 84 | 1299 | active |
| ictnlp/StreamSpeech StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p… | 37 | 1287 | active |
| studio-dots-ai/dots.tts dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive text-to-speech system, distributed as a Python library with pretrain… | 79 | 1275 | active |
| Renumics/spotlight Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe… | 93 | 1272 | active |
| ABexit/ASR-LLM-TTS An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) … | 60 | 1271 | active |
| Fictionarry/ER-NeRF ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr… | 24 | 1260 | stable |
| 0xHJK/music-dl A Python 3 command line tool that aggregates search across multiple Chinese music sites (NetEase, QQ Music, Kugou, Baidu, Xiami, Migu) and … | 39 | 4436 | maintenance |
| ryokun6/ryos ryOS is a web-based desktop environment that recreates classic macOS and Windows interfaces in the browser, built with React and TypeScript… | 84 | 1239 | active |
| DroppedNeedle/DroppedNeedle DroppedNeedle is a self-hosted music request and discovery application with a built-in library and download engine that replaces Lidarr. It… | 85 | 1229 | active |
| Tencent-Hunyuan/HunyuanCustom HunyuanCustom is a multimodal-driven customized video generation framework built on HunyuanVideo, supporting image, text, audio, and video … | 40 | 1227 | active |
| yeyupiaoling/Whisper-Finetune A toolkit for fine-tuning OpenAI's Whisper speech recognition models using LoRA, supporting training with or without timestamps and even wi… | 66 | 1223 | active |
| Aratako/Irodori-TTS Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer… | 58 | 1219 | active |
| liamcottle/reticulum-meshchat Reticulum MeshChat is an Electron-based mesh network communications app built on the Reticulum Network Stack. It lets users exchange encryp… | 88 | 1218 | active |
| EasyRPG/Player EasyRPG Player is a free, cross-platform game interpreter that plays games created with RPG Maker 2000 and 2003, using the liblcf library t… | 74 | 1217 | active |
| modal-labs/quillman QuiLLMan is a voice chat application built on Kyutai's Moshi speech-to-speech language model, deployed serverlessly on Modal with a FastAPI… | 68 | 1213 | active |
| hkjarral/AVA-AI-Voice-Agent-for-Asterisk An open-source AI voice agent that integrates with Asterisk/FreePBX phone systems via Audiosocket/RTP, built in Python with a modular pipel… | 85 | 1202 | active |
| kizuna-ai-lab/sokuji Sokuji is a cross-platform real-time two-way speech translation app for bilingual meetings, available as a desktop application (Windows, ma… | 83 | 1199 | active |
| alyssaxuu/motionity Motionity is a free, open-source, web-based motion graphics and animation editor, combining features of After Effects and Canva. It support… | 32 | 4095 | maintenance |
| BnanZ0/ok-nte ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated… | 78 | 1174 | active |
| nari-labs/dia2 Dia2 is a streaming dialogue text-to-speech model from Nari Labs that can generate conversational audio in realtime without needing the ful… | 41 | 1172 | active |
| wladradchenko/wunjo.wladradchenko.ru Wunjo CE is an open-source, locally-run AI media suite for face swap, lip sync, voice cloning, object/text/background removal, restyling, a… | 70 | 1169 | active |
| N0VI028/JS-Slash-Runner Tavern-Helper (JS-Slash-Runner) is a SillyTavern extension that lets users run external JavaScript inside the chat via sandboxed iframes. I… | 67 | 1164 | active |
| cocos2d/cocos2d-objc Cocos2D-ObjC is a 2D game engine framework for building games and graphical interactive applications on iOS, Mac, and tvOS using Objective-… | 23 | 4041 | maintenance |
| egret-labs/egret-core Egret is an open-source HTML5 game engine written in TypeScript that provides 2D and 3D rendering, GUI systems, audio, and resource managem… | 23 | 4018 | maintenance |
| TensorSpeech/TensorFlowTTS TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me… | 23 | 3995 | maintenance |
| Woolverine94/biniou biniou is a self-hosted web UI for 30+ generative AI models covering image, video, audio, and text generation, built with Gradio and Huggin… | 63 | 1150 | active |
| astronautlevel2/Anemone3DS Anemone3DS is a theme and boot splash screen manager for the Nintendo 3DS console, written in C as a homebrew application. It lets users in… | 67 | 1142 | active |
| Shirakumo/trial Trial is a modular game engine written in Common Lisp, built around OpenGL rendering with subsystems for animation, physics, audio, assets,… | 10 | 1137 | active |
| video-db/call.md Call.md is an open-source Electron desktop app that records meetings locally, transcribes them in real time with speaker separation, and pr… | 59 | 1132 | active |
| MagicFoundation/Alcinoe Alcinoe is a library of components and utilities for Delphi/FireMonkey (FMX) that helps developers build fast, modern, cross-platform appli… | 91 | 1125 | active |
| PromtEngineer/Verbi Verbi is a modular Python voice assistant application for experimenting with state-of-the-art transcription, LLM response generation, and t… | 48 | 1124 | active |
| Xilinx/Vitis_Libraries A collection of open-source, performance-optimized accelerated libraries for the AMD Vitis Unified Software Platform, offering out-of-the-b… | 87 | 1122 | active |
| HITsz-TMG/Uni-MoE Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,… | 68 | 1116 | active |
| Escartem/AnimeStudio A Unity game asset extraction tool that opens Unity bundles and lets users browse, preview, and export textures, meshes, animations, audio,… | 64 | 1112 | active |
| KoljaB/RealtimeVoiceChat A Python client-server application enabling natural spoken conversations with LLMs, streaming browser audio via WebSockets through Realtime… | 33 | 3832 | maintenance |
| vocodedev/vocode-core Vocode is an open-source Python library for building real-time, voice-based LLM applications and agents. It orchestrates streaming transcri… | 21 | 3785 | maintenance |
| doxas/twigl twigl.app is an online editor for writing ultra-compact 'one tweet' GLSL shaders for WebGL and WebGL2, with modes that auto-complete boiler… | 64 | 1085 | active |
| ARM-software/CMSIS-DSP CMSIS-DSP is ARM's optimized embedded compute library providing DSP and math kernels for Cortex-M and Cortex-A processors, including FFT, f… | 96 | 1073 | stable |
| AILab-CVC/UniRepLKNet UniRepLKNet is a large-kernel ConvNet architecture (CVPR 2024, TPAMI 2025) that provides universal perception across image, audio, video, p… | 43 | 1072 | stable |
| ufal/whisper_streaming A Python library that turns Whisper-like speech recognition models into a real-time streaming transcription and translation system using a … | 53 | 3672 | maintenance |
| brenpoly/be-more-agent An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp… | 64 | 1060 | active |
| C-Loftus/QuickPiperAudiobook A Go CLI tool that converts text content from formats like epub, PDF, mobi, txt, HTML, and docx into natural-sounding audiobooks with a sin… | 49 | 1057 | active |
| X-LANCE/SLAM-LLM SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc… | 55 | 1056 | active |
| craftyjs/Crafty Crafty is an open-source JavaScript/HTML5 game engine built around an entity-component-system architecture. It provides components for 2D g… | 23 | 3602 | maintenance |
| xiaochong/hi-kid HiKid is a free, open-source Electron desktop app that lets children in non-English-speaking countries practice English speaking and listen… | 54 | 1043 | active |
| PatterAI/Patter Patter is an open-source, MIT-licensed SDK (Python and TypeScript) that connects AI agents to real phone calls, handling telephony, speech-… | 78 | 1039 | active |
| vastxie/99AI 99AI is a commercially viable, self-hostable AI web platform built with Vue and Node.js that bundles AI chat, image/video/music generation,… | 35 | 1029 | active |
| Goekdeniz-Guelmez/Local-NotebookLM A local, open-source alternative to Google's NotebookLM that converts PDF documents into audio content like podcasts, summaries, and interv… | 74 | 1027 | active |
| EchoMimic EchoMimic is a series of open-source models (V1-V3) from Ant Group for audio-driven human animation, generating lifelike talking-head, port… | 50 | 1027 | active |