function: audio-processing
1675 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| guohuiyuan/go-music-dl A Go-based music search and download tool aggregating 10+ platforms (NetEase, QQ Music, Kugou, Bilibili, etc.) with lossless FLAC support, … | 82 | 4035 | active |
| AzuraCast/AzuraCast AzuraCast is a self-hosted, all-in-one web radio management suite that packages a full open-source radio software stack (Icecast, Liquidsoa… | 67 | 4007 | active |
| hanshuaikang/AI-Media2Doc A self-hostable web application that uses AI large language models to convert video and audio into various document styles such as Xiaohong… | 57 | 3994 | active |
| QwenLM/Qwen3-Omni Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an… | 52 | 3980 | active |
| p2r3/convert A universal online file converter that runs conversions in the browser without uploading files to a server. It supports conversions across … | 63 | 3955 | active |
| google-research/scenic Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr… | 76 | 3821 | active |
| kaldi-asr/kaldi Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea… | 52 | 15469 | maintenance |
| NExT-GPT/NExT-GPT NExT-GPT is an end-to-end any-to-any multimodal large language model that accepts and generates arbitrary combinations of text, image, vide… | 37 | 3638 | active |
| Haivision/srt SRT (Secure Reliable Transport) is an open-source transport protocol and C++ library for ultra-low-latency live video and audio streaming o… | 95 | 3585 | stable |
| danog/MadelineProto MadelineProto is an async PHP library implementing the Telegram MTProto protocol, allowing full client functionality (user or bot login) wi… | 92 | 3501 | active |
| HeapsIO/heaps Heaps is a high-performance, cross-platform 2D and 3D game engine and graphics framework written in Haxe, created by the designer of the Ha… | 74 | 3499 | stable |
| Ajaxy/telegram-tt Telegram Web A, an official Telegram web client built with a custom React-like framework (Teact) and GramJS MTProto implementation. It is a… | 98 | 3496 | active |
| huangjunsen0406/py-xiaozhi py-xiaozhi is an open-source, cross-platform multimodal AI voice assistant client written in Python, compatible with the xiaozhi-esp32 ecos… | 85 | 3452 | active |
| Kedreamix/Linly-Dubbing Linly-Dubbing is an intelligent multi-language AI dubbing and video translation tool that combines speech recognition (WhisperX, FunASR), L… | 27 | 3331 | active |
| matrixcascade/PainterEngine PainterEngine is a cross-platform application/game engine written in C with a built-in software renderer, portable to any platform supporti… | 70 | 3312 | active |
| XiaoMi/xiaomi-miloco Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an… | 82 | 3292 | active |
| OvenMediaLabs/OvenMediaEngine OvenMediaEngine is an open-source (AGPL-3.0) C++ live streaming server that ingests streams via WebRTC, SRT, RTMP, RTSP, and MPEG-2 TS and … | 90 | 3266 | active |
| MisoLabsAI/MisoTTS Miso TTS 8B is an open-source text-to-speech model based on an RVQ Transformer architecture with a Llama 3.2-style 8B backbone, designed fo… | 52 | 3224 | active |
| gozfree/gear-lib Gear-Lib is a collection of POSIX C libraries for IoT, embedded, and network service development, covering data structures, networking prot… | 23 | 3223 | active |
| zai-org/GLM-4-Voice GLM-4-Voice is an end-to-end bilingual (Chinese/English) speech dialogue model from Zhipu AI, built on GLM-4-9B with a speech tokenizer and… | 22 | 3222 | active |
| wendy7756/AI-Video-Transcriber An open-source AI tool that transcribes, summarizes, and archives videos and podcasts from 30+ platforms (YouTube, TikTok, Bilibili, etc.) … | 62 | 3215 | active |
| gcui-art/suno-api An unofficial open-source API wrapper for Suno.ai's AI music generation service, exposing endpoints for generating songs, lyrics, and audio… | 56 | 3184 | active |
| hydrusnetwork/hydrus Hydrus Network is a desktop media management application that organizes large file collections with tags instead of folders, in a booru-sty… | 95 | 3175 | active |
| facebookresearch/tribev2 TRIBE v2 is a multimodal deep learning model from Meta AI that predicts fMRI brain responses to naturalistic video, audio, and text stimuli… | 55 | 3172 | active |
| AIEraDev/Clypra Clypra is a free, open-source desktop and mobile video editor built with Tauri v2, React 19, and a Rust/FFmpeg backend. It offers multi-tra… | 81 | 3141 | active |
| matthartman/ghost-pepper A free, open-source macOS menu bar app providing fully on-device speech-to-text dictation and meeting transcription using local Whisper, Pa… | 78 | 3140 | active |
| g3n/engine G3N is an OpenGL-based 3D game engine written in Go, with an integrated GUI framework and 3D spatial audio via OpenAL. It can be used to bu… | 66 | 3110 | active |
| 4sval/FModel FModel is a desktop archive explorer for Unreal Engine game packages (PAK files), built in C# on top of the CUE4Parse library. It provides … | 87 | 3090 | active |
| miroslavpejic85/mirotalksfu MiroTalk SFU is a self-hosted, open-source WebRTC video conferencing platform built on the Mediasoup SFU architecture, positioning itself a… | 77 | 3087 | active |
| elevenlabs/elevenlabs-python The official Python SDK for the ElevenLabs API, providing programmatic access to text-to-speech, speech-to-text, voice cloning, dubbing, mu… | 93 | 3078 | active |
| korlibs/korge KorGE is a modern multiplatform game engine written entirely in Kotlin, built on top of the Korlibs multimedia stack. It targets JVM/Androi… | 67 | 3041 | active |
| q191201771/lal LAL is an audio/video live streaming broadcast server written in Go, comparable to nginx-rtmp-module but with more features. It supports RT… | 23 | 3026 | active |
| MeiGen-AI/MultiTalk MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a … | 56 | 2992 | active |
| TeamUltroid/Ultroid Ultroid is a multi-featured Telegram userbot built in Python on the Telethon library, with a pluggable architecture and voice/video call mu… | 65 | 2976 | active |
| GaijinEntertainment/DagorEngine Dagor Engine is the proprietary game engine and toolchain source code from Gaijin Games (creators of War Thunder), released for public use.… | 89 | 2953 | active |
| InternLM/InternLM-XComposer InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u… | 38 | 2925 | active |
| state-spaces/s4 Official implementations of Structured State Space (S4) sequence models and related variants like S4D, HiPPO, and SaShiMi, built in PyTorch… | 32 | 2919 | active |
| datascale-ai/opentalking OpenTalking is an open-source Python framework for building real-time AI digital-human (talking avatar) conversation products. It orchestra… | 65 | 2897 | active |
| quickshell-mirror/quickshell Quickshell is a toolkit for building desktop shell components such as status bars, widgets, lockscreens, and display managers using QtQuick… | 83 | 2873 | active |
| zig-gamedev/zig-gamedev The main development repository for the zig-gamedev organization, containing Zig game development libraries and a large collection of sampl… | 64 | 2862 | active |
| jhj0517/Whisper-WebUI A Gradio-based web interface for OpenAI's Whisper models that generates subtitles from files, YouTube videos, or microphone input. It suppo… | 63 | 2862 | active |
| microsoft/DirectXTK The DirectX Tool Kit (DirectXTK) is a collection of helper classes for writing Direct3D 11 C++ code in Win32 desktop, UWP, and Xbox One app… | 85 | 2851 | stable |
| bytedeco/javacpp-presets JavaCPP Presets provides Java bindings for commonly used native C++ libraries such as OpenCV, FFmpeg, and many others, built on the JavaCPP… | 86 | 2850 | active |
| unknownskl/greenlight Greenlight is an open-source desktop client for streaming games from Xbox Cloud Gaming (xCloud) and Xbox consoles, built in TypeScript as a… | 87 | 2844 | active |
| scummvm/scummvm ScummVM is a cross-platform application that lets you play classic point-and-click adventure games, text adventures, and RPGs by replacing … | 87 | 2792 | stable |
| AutoArk/GPA GPA (General Purpose Audio) is a unified autoregressive audio-language model that performs text-to-speech, automatic speech recognition, an… | 54 | 2762 | active |
| apple/turicreate Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj… | 10 | 11159 | maintenance |
| rhasspy/rhasspy Rhasspy is a fully offline, privacy-focused set of voice assistant services supporting many human languages. It converts spoken voice comma… | 10 | 2750 | active |
| FWGS/xash3d-fwgs Xash3D FWGS is a cross-platform game engine forked from Xash3D, aimed at compatibility with the Half-Life (GoldSrc) engine while extending … | 86 | 2739 | active |
| Suxiaoqinx/Netease_url A self-hosted Python service that parses Netease Cloud Music (网易云音乐) songs, playlists, and albums to retrieve direct download URLs at multi… | 73 | 2673 | active |
| Const-me/Whisper A Windows port of whisper.cpp that runs OpenAI's Whisper speech recognition model on the GPU via DirectCompute (Direct3D 11 compute shaders… | 60 | 10649 | maintenance |
| morethanwords/tweb Telegram Web K is the open-source TypeScript web client powering web.telegram.org/k/, based on the original Webogram and actively patched a… | 77 | 2623 | active |
| bjornbytes/lovr LÖVR is an open-source Lua framework for rapidly building 3D games and VR experiences, written in C11 and scripted with LuaJIT. It provides… | 83 | 2587 | active |
| ahmedeltaher/Android-MVVM-Architecture-Android-Voice-AI-SDK A reusable Android library (Kotlin, MVVM) that provides a full voice-driven AI conversation pipeline: microphone capture with VAD, speech-t… | 70 | 2578 | active |
| mozilla/TTS A deep learning library for advanced text-to-speech generation, built on PyTorch with models like Tacotron2, Glow-TTS, and various vocoders… | 23 | 10167 | maintenance |
| get-iplayer/get_iplayer A Perl-based command-line tool that indexes and downloads TV and radio programmes from BBC iPlayer and BBC Sounds. It supports regex search… | 31 | 2538 | active |
| ChangbaDevs/KTVHTTPCache KTVHTTPCache is an Objective-C HTTP caching framework for iOS that caches multimedia resources by acting as a local proxy server, enabling … | 50 | 2508 | stable |
| janhq/ichigo Ichigo is a Python speech package for developers offering local realtime voice AI capabilities, including a compact 22M-parameter speech to… | 50 | 2492 | active |
| dino/dino Dino is a modern open-source XMPP (Jabber) chat client for Linux desktops, built with GTK4 and Vala. It supports end-to-end encryption via … | 76 | 2480 | active |
| erew123/alltalk_tts AllTalk TTS is a text-to-speech application built on the Coqui TTS engine, usable standalone or as an extension for Text-generation-webui, … | 44 | 2429 | active |
| pnnbao97/VieNeu-TTS VieNeu-TTS is an on-device Vietnamese text-to-speech library with instant zero-shot voice cloning from short reference clips, supporting bi… | 84 | 2427 | active |
| nesrak1/UABEA UABEA is a cross-platform Unity asset bundle and serialized file reader/writer built in C# with Avalonia. It is designed as a modding and r… | 59 | 2404 | active |
| ValveResourceFormat/ValveResourceFormat Source 2 Viewer (VRF) is an open-source tool for browsing VPK archives and viewing, extracting, and decompiling Source 2 game assets such a… | 95 | 2391 | active |
| ailia-ai/ailia-models A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,… | 77 | 2385 | active |
| vercel/ai-elements AI Elements is a React component library and custom shadcn/ui registry from Vercel providing pre-built, composable UI components for AI-nat… | 75 | 2365 | active |
| snapotter-hq/SnapOtter SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities… | 80 | 2362 | active |
| elevenlabs/ui ElevenLabs UI is a component library and custom shadcn/ui registry providing pre-built React components for building multimodal, audio, and… | 56 | 2361 | active |
| facebookresearch/perception_models Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan… | 54 | 2353 | active |
| Xilinx/PYNQ PYNQ is an open-source Python framework from AMD/Xilinx for designing embedded systems on Zynq and other adaptive computing platforms (FPGA… | 78 | 2339 | active |
| stephengpope/no-code-architects-toolkit A self-hostable Flask-based API that consolidates common media processing tasks—video editing, captioning, audio conversion, transcription,… | 50 | 2339 | active |
| jeremyckahn/chitchatter Chitchatter is a free, open-source communication tool offering secure peer-to-peer chat, video, audio, and file sharing directly between br… | 75 | 2321 | active |
| facebookresearch/ImageBind A PyTorch library from Meta AI implementing ImageBind, a model that learns a joint embedding space across six modalities: images, text, aud… | 54 | 9064 | maintenance |
| mattiasgustavsson/libs A collection of single-file, public domain (dual MIT) C/C++ libraries covering graphics app frameworks, data structures, threading, HTTP, a… | 61 | 2299 | active |
| QwenAudio/qwen-audio-agent A realtime voice runtime and frontend for AI coding agents like Claude Code, Codex, and Qwen Code, letting agents talk, listen, and report … | 79 | 2264 | active |
| uowuo/abaddon Abaddon is an alternative Discord client written in C++ using GTK 3, offering a lightweight non-Electron desktop experience with voice chat… | 73 | 2216 | active |
| Hubs-Foundation/hubs Hubs is an open-source, browser-based multi-user 3D virtual world and social VR platform built with A-Frame, Three.js, and WebXR/WebRTC. It… | 82 | 2214 | active |
| DigitalPhonetics/IMS-Toucan IMS Toucan is a PyTorch-based toolkit for training and running state-of-the-art, controllable text-to-speech synthesis, home of the massive… | 63 | 2207 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2203 | active |
| TransWithAI/Faster-Whisper-TransWithAI-ChickenRice A high-performance audio/video transcription and translation application built on Faster Whisper, optimized for Japanese-to-Chinese transla… | 77 | 2196 | active |
| meizhong986/WhisperJAV WhisperJAV is a local, privacy-preserving subtitle generator for Japanese adult videos, combining Qwen3-ASR, Whisper, TEN-VAD/FireRedVAD vo… | 94 | 2170 | active |
| openfl/openfl OpenFL is an open-source Haxe framework implementing the classic Flash/AIR APIs for building games and creative applications. A single code… | 89 | 2152 | active |
| OpenVidu/openvidu OpenVidu is a self-hosted, open-source video conferencing platform and WebRTC SDK suite, built as a LiveKit fork with mediasoup as its SFU … | 91 | 2126 | active |
| selkies-project/selkies Selkies is an open-source low-latency, GPU/CPU-accelerated Linux remote desktop and game streaming platform that delivers an HTML5 web clie… | 67 | 2067 | active |
| travisvn/openai-edge-tts A self-hosted Python service that emulates the OpenAI text-to-speech API endpoint (/v1/audio/speech) using Microsoft Edge's free online TTS… | 26 | 2066 | active |
| jaywalnut310/vits VITS is the official PyTorch implementation of an end-to-end text-to-speech model based on a conditional variational autoencoder with adver… | 32 | 7889 | maintenance |
| ossia/score ossia score is a free, open-source interactive sequencer for audio-visual artists, designed to create interactive shows, installations, and… | 92 | 2053 | active |
| 0xShug0/audio.cpp audio.cpp is a pure C++ inference engine for audio models built on ggml, supporting TTS, STT, VAD, voice conversion, music generation, and … | 80 | 2022 | active |
| ppwwyyxx/wechat-dump A Python-based tool that extracts and parses WeChat message history from a rooted Android phone, decoding the local message database and me… | 54 | 2010 | active |
| mxpv/podsync Podsync is a self-hosted service and CLI that converts YouTube, Vimeo, or Soundcloud channels, users, and playlists into podcast RSS feeds.… | 67 | 1952 | active |
| GameFoundry/B3DFramework B3D Framework (formerly Banshee 3D) is a modern C++17 library providing a foundation for high-performance game engines and real-time graphi… | 77 | 1922 | active |
| OpenTalker/video-retalking VideoReTalking is a Python research system from SIGGRAPH Asia 2022 that edits real-world talking-head videos to match a given audio track, … | 23 | 7280 | maintenance |
| diffgram/diffgram Diffgram is a self-hosted AI datastore for managing schemas, BLOBs, and predictions, with built-in human supervision (data labeling), data … | 62 | 1909 | active |
| ZhipingYang/UUChatTableView UUChatTableView is a lightweight Objective-C UIKit component that renders group and one-to-one chat interfaces with text, image, and voice … | 57 | 1906 | active |
| Quran.com The official open-source frontend for Quran.com, built with Next.js and TypeScript, providing a web application for reading, listening to, … | 65 | 1901 | active |
| lanbinleo/bili2text bili2text is a Python command-line tool that converts Bilibili videos into text transcripts from a link or BV number, handling download, au… | 56 | 1873 | active |
| LumePart/Explo Explo is a self-hosted web application that brings Spotify's Discover Weekly-style music discovery to self-hosted music servers like Navidr… | 91 | 1859 | active |
| Game Framework Game Framework is a game framework built on the Unity engine that encapsulates 19 commonly used game development modules such as entities, … | 32 | 6849 | maintenance |
| CoderLine/alphaTab alphaTab is a cross-platform music notation and guitar tablature rendering library available for JavaScript, .NET, and Android. It loads sc… | 92 | 1809 | active |
| yerfor/GeneFacePlusPlus GeneFace++ is the official PyTorch implementation of a NeRF-based system for generalized and stable real-time 3D talking face generation. I… | 26 | 1809 | active |
| whotto/Video_note_generator A Python tool that converts video URLs into polished Xiaohongshu (Little Red Book) notes and blog articles. It downloads videos, transcribe… | 43 | 1808 | active |