Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: audio-processing

1675 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
SunoAI-API/Suno-API
An unofficial Python/FastAPI-based API wrapper for Suno AI that lets you generate songs and lyrics programmatically. It includes automatic …
331806active
FL33TW00D/whisper-turbo
Whisper Turbo is a fast, cross-platform, GPU-accelerated implementation of OpenAI's Whisper speech recognition model, built on the Ratchet …
191794active
HFrost0/bilix
bilix is a lightning-fast asynchronous download tool for bilibili and other sites, usable via a CLI or as an extensible Python library. It …
231787active
wwbin2017/bailing
Bailing is an open-source voice assistant application similar to GPT-4o, built with an ASR + VAD + LLM + TTS pipeline (FunASR, silero-vad, …
541757active
microsoft/DirectXTK12
DirectX Tool Kit for DirectX 12 is a C++ library of helper classes for writing Direct3D 12 code, covering sprite rendering, effects, textur…
851749stable
elder-plinius/ST3GG
ST3GG is an all-in-one steganography toolkit that hides secret data inside images, audio, documents, and network packets using 100+ encodin…
631741active
NVIDIA-NeMo/Curator
NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for …
861736active
AutoArk/EVA-OS
EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining …
681712active
fdivitto/FabGL
FabGL is a C++ graphics and input library for the ESP32, providing display drivers for VGA, PAL/NTSC composite video, and I2C/SPI LCDs, plu…
261703active
OpenMOSS/MOSS-Transcribe-Diarize
MOSS-Transcribe-Diarize 0.9B is an open-source end-to-end audio understanding model that jointly performs multi-speaker speech transcriptio…
581695active
oakmound/oak
Oak is a pure Go 2D game engine providing rendering, particles, collision, physics, audio, input, and event systems without external depend…
661670active
TorqueGameEngines/Torque2D
Torque2D is a free, MIT-licensed, open-source 2D game engine in C++ built on the proven GarageGames Torque technology. It offers cross-plat…
901667active
Fantasy-AMAP/fantasy-talking
FantasyTalking is a research codebase and model for generating realistic talking portrait videos from a single image and an audio clip, bui…
481628active
ZiqiaoPeng/SyncTalk
SyncTalk is the official PyTorch implementation of a CVPR 2024 paper that synthesizes speech-driven, synchronized talking head videos using…
461626active
JohnEarnest/Decker
Decker is a multimedia platform and sketchpad for creating and sharing interactive documents with sound, images, hypertext, and scripted be…
761615active
hiloteam/Hilo
Hilo is a cross-end HTML5 game development framework from Alibaba Group, providing a lightweight visual object architecture with multiple r…
445939maintenance
ml4a/ml4a
ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li…
321602active
X-LANCE/AniTalker
AniTalker is the official PyTorch implementation of an ACM MM 2024 paper that animates a single static portrait into a vivid talking-face v…
241598active
akudamatata/Solara
Solara is a minimalist web-based music player that aggregates multiple free music APIs for search, streaming playback, and audio downloads.…
571582active
songloft-org/songloft
Songloft (formerly MiMusic) is a free, ad-free, plugin-based self-hosted music server written in Go, designed for managing music you legall…
821562active
jpush/aurora-imui
Aurora IMUI is a general-purpose instant messaging UI component library providing MessageList and InputView components, independent of any …
235696maintenance
shrimbly/node-banana
Node Banana is an open-source, node-based visual workflow editor for building AI media generation pipelines. Users connect nodes on an infi…
811550active
sepfy/libpeer
libpeer is a portable WebRTC implementation written in C using BSD sockets, designed for IoT and embedded devices such as ESP32 and Raspber…
771542active
vibevoice-community/VibeVoice
VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker …
601542active
tin2tin/Pallaidium
Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE…
751520active
Flashlight
wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into …
625466maintenance
Taiizor/Sucrose
Sucrose is a free, open-source wallpaper engine for Windows that renders interactive live wallpapers from GIFs, videos, URLs, web pages, Yo…
921503active
torch2424/wasmboy
WasmBoy is a Game Boy and Game Boy Color emulator library written in WebAssembly using AssemblyScript, distributed via npm and importable i…
321498active
QwenAudio/Fun-ASR
Fun-ASR is a family of open-source LLM-based end-to-end speech recognition models from Tongyi Lab, covering Chinese, dialects, accents, and…
811496active
p2r3/beheader
A command-line tool that generates polyglot files - single files that are simultaneously valid images, videos, PDFs, ZIP archives, and HTML…
481455active
silverstein/minutes
Minutes is an open-source, privacy-first conversation memory app that records meetings, voice memos, and dictation, transcribes them locall…
771453active
shridarpatil/whatomate
Whatomate is an open-source, self-hosted WhatsApp Business platform shipped as a single Go binary with Docker support. It provides multi-te…
691450active
voice-cloning-app/Voice-Cloning-App
A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom…
231440active
devnen/Chatterbox-TTS-Server
A self-hosted server wrapping Resemble AI's Chatterbox TTS models behind an OpenAI-compatible API with a modern web UI. It supports voice c…
651420active
K0lb3/UnityPy
UnityPy is a Python library for extracting, unpacking, and editing Unity game engine asset files, based on AssetStudio. It supports reading…
751414active
mpiorowski/late-sh
late.sh is a terminal-first social platform accessed entirely over SSH, offering real-time chat, lofi/classical radio streaming, terminal g…
771408active
Arthi-chaud/Meelo
Meelo is a self-hosted music streaming server designed for music collectors, similar to Plex or Jellyfin but focused on music. It offers ri…
991398active
gerbera/gerbera
Gerbera is a free, open-source UPnP/DLNA media server that streams digital media across a home network to compatible devices like TVs, game…
861390active
lamm-mit/PDF2Audio
A Gradio-based web application that converts PDF documents into audio podcasts, lectures, and summaries using OpenAI GPT models for text ge…
301383active
brainboxdotcc/DPP
D++ (DPP) is a lightweight, modern C++ library for the Discord API, supporting API v10 with efficient caching, sharding, clustering, slash …
981379stable
zelon88/HRConvert2
HRConvert2 is a self-hosted, resource-aware file conversion server written in PHP that supports 488 file formats across documents, images, …
991364active
elevenyellow/handcrafted-persona-engine
Persona Engine is a Windows desktop application that drives a Live2D avatar with an AI pipeline: microphone speech recognition, an LLM guid…
731357active
yuanyuanxiang/SimpleRemoter
SimpleRemoter (YAMA) is a C++ remote control suite derived from the Gh0st RAT codebase, providing remote desktop, file transfer, terminal, …
601347active
AI-FanGe/OpenAIglasses_for_Navigation
An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st…
391343active
numandev1/react-native-compressor
A React Native library that compresses images, videos, and audio with WhatsApp-like quality, plus background upload, file download, and vid…
961325active
copperspice/copperspice
CopperSpice is a set of cross-platform C++ libraries (Core, Gui, Network, Multimedia, SQL, OpenGL, Vulkan, WebKit, XML, and more) derived f…
741324active
caiiiycuk/js-dos
js-dos is a TypeScript library and API for running DOS and Windows 9x programs in the browser or Node.js, acting as a frontend for DOSBox/D…
971322active
STMicroelectronics/STM32CubeF4
STMicroelectronics' official STM32Cube firmware package for the STM32F4 series of microcontrollers. It bundles HAL and LL drivers, CMSIS co…
721316stable
DAMO-NLP-SG/VideoLLaMA2
VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro…
251307active
StarlightSearch/EmbedAnything
EmbedAnything is a high-performance, memory-safe embedding pipeline written in Rust (with Python bindings) that generates embeddings from t…
881305active
Henry-23/VideoChat
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late…
481303active
stenolabs/stenoai
Steno is a privacy-first desktop AI notepad and meeting notetaker that records, transcribes, summarizes, and lets you query meetings entire…
841299active
ictnlp/StreamSpeech
StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, p…
371287active
studio-dots-ai/dots.tts
dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive text-to-speech system, distributed as a Python library with pretrain…
791275active
Renumics/spotlight
Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe…
931272active
ABexit/ASR-LLM-TTS
An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) …
601271active
Fictionarry/ER-NeRF
ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr…
241260stable
0xHJK/music-dl
A Python 3 command line tool that aggregates search across multiple Chinese music sites (NetEase, QQ Music, Kugou, Baidu, Xiami, Migu) and …
394436maintenance
ryokun6/ryos
ryOS is a web-based desktop environment that recreates classic macOS and Windows interfaces in the browser, built with React and TypeScript…
841239active
DroppedNeedle/DroppedNeedle
DroppedNeedle is a self-hosted music request and discovery application with a built-in library and download engine that replaces Lidarr. It…
851229active
Tencent-Hunyuan/HunyuanCustom
HunyuanCustom is a multimodal-driven customized video generation framework built on HunyuanVideo, supporting image, text, audio, and video …
401227active
yeyupiaoling/Whisper-Finetune
A toolkit for fine-tuning OpenAI's Whisper speech recognition models using LoRA, supporting training with or without timestamps and even wi…
661223active
Aratako/Irodori-TTS
Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer…
581219active
liamcottle/reticulum-meshchat
Reticulum MeshChat is an Electron-based mesh network communications app built on the Reticulum Network Stack. It lets users exchange encryp…
881218active
EasyRPG/Player
EasyRPG Player is a free, cross-platform game interpreter that plays games created with RPG Maker 2000 and 2003, using the liblcf library t…
741217active
modal-labs/quillman
QuiLLMan is a voice chat application built on Kyutai's Moshi speech-to-speech language model, deployed serverlessly on Modal with a FastAPI…
681213active
hkjarral/AVA-AI-Voice-Agent-for-Asterisk
An open-source AI voice agent that integrates with Asterisk/FreePBX phone systems via Audiosocket/RTP, built in Python with a modular pipel…
851202active
kizuna-ai-lab/sokuji
Sokuji is a cross-platform real-time two-way speech translation app for bilingual meetings, available as a desktop application (Windows, ma…
831199active
alyssaxuu/motionity
Motionity is a free, open-source, web-based motion graphics and animation editor, combining features of After Effects and Canva. It support…
324095maintenance
BnanZ0/ok-nte
ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated…
781174active
nari-labs/dia2
Dia2 is a streaming dialogue text-to-speech model from Nari Labs that can generate conversational audio in realtime without needing the ful…
411172active
wladradchenko/wunjo.wladradchenko.ru
Wunjo CE is an open-source, locally-run AI media suite for face swap, lip sync, voice cloning, object/text/background removal, restyling, a…
701169active
N0VI028/JS-Slash-Runner
Tavern-Helper (JS-Slash-Runner) is a SillyTavern extension that lets users run external JavaScript inside the chat via sandboxed iframes. I…
671164active
cocos2d/cocos2d-objc
Cocos2D-ObjC is a 2D game engine framework for building games and graphical interactive applications on iOS, Mac, and tvOS using Objective-…
234041maintenance
egret-labs/egret-core
Egret is an open-source HTML5 game engine written in TypeScript that provides 2D and 3D rendering, GUI systems, audio, and resource managem…
234018maintenance
TensorSpeech/TensorFlowTTS
TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me…
233995maintenance
Woolverine94/biniou
biniou is a self-hosted web UI for 30+ generative AI models covering image, video, audio, and text generation, built with Gradio and Huggin…
631150active
astronautlevel2/Anemone3DS
Anemone3DS is a theme and boot splash screen manager for the Nintendo 3DS console, written in C as a homebrew application. It lets users in…
671142active
Shirakumo/trial
Trial is a modular game engine written in Common Lisp, built around OpenGL rendering with subsystems for animation, physics, audio, assets,…
101137active
video-db/call.md
Call.md is an open-source Electron desktop app that records meetings locally, transcribes them in real time with speaker separation, and pr…
591132active
MagicFoundation/Alcinoe
Alcinoe is a library of components and utilities for Delphi/FireMonkey (FMX) that helps developers build fast, modern, cross-platform appli…
911125active
PromtEngineer/Verbi
Verbi is a modular Python voice assistant application for experimenting with state-of-the-art transcription, LLM response generation, and t…
481124active
Xilinx/Vitis_Libraries
A collection of open-source, performance-optimized accelerated libraries for the AMD Vitis Unified Software Platform, offering out-of-the-b…
871122active
HITsz-TMG/Uni-MoE
Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,…
681116active
Escartem/AnimeStudio
A Unity game asset extraction tool that opens Unity bundles and lets users browse, preview, and export textures, meshes, animations, audio,…
641112active
KoljaB/RealtimeVoiceChat
A Python client-server application enabling natural spoken conversations with LLMs, streaming browser audio via WebSockets through Realtime…
333832maintenance
vocodedev/vocode-core
Vocode is an open-source Python library for building real-time, voice-based LLM applications and agents. It orchestrates streaming transcri…
213785maintenance
doxas/twigl
twigl.app is an online editor for writing ultra-compact 'one tweet' GLSL shaders for WebGL and WebGL2, with modes that auto-complete boiler…
641085active
ARM-software/CMSIS-DSP
CMSIS-DSP is ARM's optimized embedded compute library providing DSP and math kernels for Cortex-M and Cortex-A processors, including FFT, f…
961073stable
AILab-CVC/UniRepLKNet
UniRepLKNet is a large-kernel ConvNet architecture (CVPR 2024, TPAMI 2025) that provides universal perception across image, audio, video, p…
431072stable
ufal/whisper_streaming
A Python library that turns Whisper-like speech recognition models into a real-time streaming transcription and translation system using a …
533672maintenance
brenpoly/be-more-agent
An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp…
641060active
C-Loftus/QuickPiperAudiobook
A Go CLI tool that converts text content from formats like epub, PDF, mobi, txt, HTML, and docx into natural-sounding audiobooks with a sin…
491057active
X-LANCE/SLAM-LLM
SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc…
551056active
craftyjs/Crafty
Crafty is an open-source JavaScript/HTML5 game engine built around an entity-component-system architecture. It provides components for 2D g…
233602maintenance
xiaochong/hi-kid
HiKid is a free, open-source Electron desktop app that lets children in non-English-speaking countries practice English speaking and listen…
541043active
PatterAI/Patter
Patter is an open-source, MIT-licensed SDK (Python and TypeScript) that connects AI agents to real phone calls, handling telephony, speech-…
781039active
vastxie/99AI
99AI is a commercially viable, self-hostable AI web platform built with Vue and Node.js that bundles AI chat, image/video/music generation,…
351029active
Goekdeniz-Guelmez/Local-NotebookLM
A local, open-source alternative to Google's NotebookLM that converts PDF documents into audio content like podcasts, summaries, and interv…
741027active
EchoMimic
EchoMimic is a series of open-source models (V1-V3) from Ant Group for audio-driven human animation, generating lifelike talking-head, port…
501027active

← prev page 16 / 17 next →