function: audio-processing
1675 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| rayenfeng/riko_project Project Riko is an anime-themed conversational voice assistant that combines OpenAI's GPT for dialogue, GPT-SoVITS for voice synthesis, and… | 31 | 1023 | active |
| towhee-io/towhee Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s… | 23 | 3454 | maintenance |
| pingostack/pingos PingOS is an NGINX-based streaming media server supporting RTMP, HTTP-FLV, HTTP-TS, HLS, HLS+, and DASH live streaming protocols with H.264… | 23 | 1013 | active |
| GitLqr/LQRWeChat An open-source Android app that closely replicates WeChat 6.5.7, built on the RongCloud IM SDK with RxJava, Retrofit, MVP, and Glide. It su… | 73 | 3426 | maintenance |
| TensorSpeech/TensorFlowASR TensorFlowASR is a Python library implementing automatic speech recognition architectures such as DeepSpeech2, Jasper, RNN Transducer, Cont… | 66 | 1010 | active |
| siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro… | 39 | 1001 | active |
| DAMO-NLP-SG/Video-LLaMA Video-LLaMA is an instruction-tuned audio-visual language model that extends LLaMA with video and audio understanding via cross-modal pretr… | 29 | 3139 | maintenance |
| lynckia/licode Licode is an open-source WebRTC communications platform for hosting your own videoconference provider, built around the C++ Erizo MCU, a No… | 67 | 3134 | maintenance |
| StarRTC StarRTC is a free cross-platform real-time communication SDK and self-hostable server suite providing instant messaging, one-to-one video c… | 23 | 3075 | maintenance |
| CainKernel/CainCamera An open-source Android app and set of libraries demonstrating how to build a beauty camera, image editor, and short-video editor. It implem… | 32 | 2962 | maintenance |
| yerfor/GeneFace GeneFace is the official PyTorch implementation of an ICLR 2023 paper on generalized, high-fidelity audio-driven 3D talking face synthesis … | 21 | 2657 | maintenance |
| mravanelli/pytorch-kaldi PyTorch-Kaldi is a toolkit for developing state-of-the-art DNN/HMM hybrid speech recognition systems, combining PyTorch-managed neural netw… | 32 | 2404 | maintenance |
| ivansafrin/Polycode Polycode is a cross-platform C++ framework for creative coding that provides accelerated 2D and 3D graphics, hardware shaders, sound, netwo… | 23 | 2382 | maintenance |
| fatchord/WaveRNN A PyTorch implementation of DeepMind's WaveRNN neural vocoder plus a Tacotron text-to-speech system, trained on LJSpeech. It supports train… | 32 | 2190 | maintenance |
| kisence-mian/MyUnityFrameWork A pluggable Unity game development framework written in C# covering resource loading, config/data loading, UI management, audio, logging, a… | 32 | 2189 | maintenance |
| jadepeng/XMusicDownloader A C# desktop application that aggregates music search across Baidu, NetEase, QQ, Kugou, and Migu music sites and supports batch downloading… | 23 | 2140 | maintenance |
| SeanNaren/deepspeech.pytorch A PyTorch implementation of the DeepSpeech2 speech recognition model, built on PyTorch Lightning, supporting training, testing, and inferen… | 23 | 2136 | maintenance |
| WangShuo1143368701/WSLiveDemo An Android live-streaming SDK providing RTMP pushing, video recording, and GPU-based filters including adjustable beauty effects. It uses O… | 23 | 2056 | maintenance |
| vsitzmann/siren Official PyTorch implementation of SIREN, a neural network architecture using periodic (sine) activation functions for implicit neural repr… | 32 | 1993 | maintenance |
| orangeduck/Corange Corange is a pure C game engine built on SDL and OpenGL, providing asset, entity, and UI management plus a modern deferred renderer. It ser… | 32 | 1987 | maintenance |
| nobody132/masr MASR is an end-to-end Mandarin Chinese automatic speech recognition project built on a gated convolutional neural network (similar to Wav2L… | 23 | 1968 | maintenance |
| Kyubyong/tacotron A heavily documented TensorFlow implementation of Tacotron, a fully end-to-end text-to-speech synthesis model. It includes training, prepro… | 32 | 1832 | maintenance |
| iqiqiya/iqiqiya-API A PHP-based collection of free web API endpoints for parsing media from Chinese platforms (Douyin, Kuaishou, Bilibili, NetEase Music, Ximal… | 10 | 1737 | maintenance |
| Lightning-Universe/lightning-flash Lightning Flash is a high-level PyTorch library built on PyTorch Lightning that provides ready-made 'recipes' for over 15 AI tasks across 7… | 10 | 1722 | maintenance |
| GarageGames/Torque2D Torque 2D is an MIT-licensed open source C++ 2D game engine originally from GarageGames, featuring OpenGL batched rendering, Box2D physics,… | 23 | 1663 | maintenance |
| invictus717/MetaTransformer Meta-Transformer is a research framework for unified multimodal learning that maps inputs from 12 modalities (text, images, point clouds, a… | 19 | 1647 | maintenance |
| HumanAIGC/EMO EMO (Emote Portrait Alive) is a research codebase from Alibaba's Institute for Intelligent Computing that generates expressive talking port… | 25 | 7594 | experimental |
| supermedium/superframe Superframe is a curated collection of components for A-Frame, the web framework for building VR experiences. It includes components for ani… | 70 | 1412 | maintenance |
| innnky/emotional-vits Emotional VITS is a voice synthesis (TTS) model based on VITS that enables emotion-controllable speech generation without requiring manual … | 32 | 1392 | maintenance |
| guangqiang-liu/OneM OneM is a comprehensive React Native app combining magazine browsing, music playback, and video playback, built with Redux state management… | 32 | 1375 | maintenance |
| CSTR-Edinburgh/merlin Merlin is a toolkit from the University of Edinburgh's CSTR for building deep neural network models for statistical parametric speech synth… | 23 | 1321 | maintenance |
| YuanxunLu/LiveSpeechPortraits A PyTorch implementation of the SIGGRAPH Asia 2021 paper 'Live Speech Portraits', which generates photorealistic personalized talking-head … | 32 | 1283 | maintenance |
| PlayVoice/vits_chinese A Chinese text-to-speech library combining VITS with BERT-based prosody embeddings and NaturalSpeech infer-loss features, supporting ONNX c… | 23 | 1227 | maintenance |
| kylestetz/slang Slang is a small audio programming language implemented entirely in JavaScript, running in the browser on top of the Web Audio API. It uses… | 32 | 1203 | maintenance |
| google/lullaby Lullaby is a collection of high-performance C++ libraries from Google for building virtual and augmented reality experiences. It provides a… | 10 | 1198 | maintenance |
| MaurycyLiebner/enve enve is an open-source 2D animation application for Linux and Windows, supporting vector and raster animation as well as sound and video fi… | 10 | 1195 | maintenance |
| AnyListen/YaVipCore A .NET Core music interface service that aggregates search and playback across Chinese music platforms such as NetEase, QQ, Kugou, Kuwo, Ba… | 23 | 1183 | maintenance |
| kpreid/shinysdr ShinySDR is a software-defined radio receiver application built on GNU Radio with a web-based user interface and a plugin system. Combined … | 32 | 1136 | maintenance |
| snuffyDev/Beatbump Beatbump is a privacy-respecting alternative frontend for YouTube Music built with SvelteKit, offering ad-free audio playback, playlist man… | 10 | 1132 | maintenance |
| MRzzm/DINet DINet is the official PyTorch implementation of an AAAI 2023 paper on realistic face visually dubbing, which deforms and inpaints mouth reg… | 32 | 1128 | maintenance |
| maelfabien/Multimodal-Emotion-Recognition A real-time multimodal emotion recognition web app built with Flask that analyzes emotions from text, audio, and video inputs using deep le… | 32 | 1089 | maintenance |
| HurTeng/StormPlane StormPlane (Desert Storm) is an open-source vertical-scrolling shoot-'em-up game for Android, inspired by Raiden and WeChat's plane game, b… | 23 | 1072 | maintenance |
| OFA-Sys/ONE-PEACE ONE-PEACE is a general multimodal representation model that jointly encodes vision, audio, and language modalities without initializing fro… | 29 | 1060 | maintenance |
| NATSpeech/NATSpeech A PyTorch framework for non-autoregressive text-to-speech (NAR-TTS), containing official implementations of PortaSpeech (NeurIPS 2021) and … | 23 | 1004 | maintenance |
| HailToDodongo/pyrite64 Pyrite64 is a visual editor and runtime game engine for creating 3D games that run on real Nintendo 64 consoles or accurate emulators, buil… | 82 | 3239 | experimental |
| enhuiz/vall-e An unofficial PyTorch implementation of the VALL-E text-to-speech audio language model, built on the EnCodec tokenizer. It provides trainin… | 31 | 2976 | experimental |
| jbilcke-hf/clapper Clapper is an open-source AI story visualization tool - a video synthesizer and sequencer for creating AI-generated cinema. Users iterate o… | 39 | 2331 | experimental |
| vanilla-wiiu/vanilla Vanilla is a software clone of the Wii U Gamepad that lets compatible devices act as a wireless second screen and controller for a Wii U co… | 80 | 2297 | experimental |
| EQMG/Acid Acid is an open-source, cross-platform 3D game engine written in modern C++17, using Vulkan as its sole graphics API (with Metal support vi… | 32 | 2019 | experimental |
| Standard-Intelligence/hertz-dev Hertz-dev is an open-source 8.5B parameter autoregressive base model for full-duplex conversational audio, released by Standard Intelligenc… | 22 | 1799 | experimental |
| hugoam/toy toy is a thin and modular C++ game engine built on top of the 'two' library, aiming to provide the simplest technology stack for making 2D … | 32 | 1594 | experimental |
| lyuchenyang/Macaw-LLM Macaw-LLM is a multi-modal language modeling framework that integrates image, video, audio, and text data, built on CLIP, Whisper, and LLaM… | 29 | 1591 | experimental |
| lucidrains/naturalspeech2-pytorch A PyTorch implementation of NaturalSpeech 2, a zero-shot text-to-speech and singing synthesizer that combines a neural audio codec with a l… | 20 | 1333 | experimental |
| magenta/magenta Magenta is a Google Brain research project and Python TensorFlow library for generating music, images, and drawings with machine learning. … | 10 | 19808 | abandoned |
| benawad/dogehouse DogeHouse is an open-source social audio platform (a Clubhouse-style voice conversation app) with an Elixir API, voice server, Next.js web … | 23 | 9023 | abandoned |
| CatchChat/Yep Yep is an open-source iOS social networking app written in Swift, themed around 'meeting geniuses' to help users find experts and learners … | 10 | 5871 | abandoned |
| amhndu/SimpleNES SimpleNES is an NES (Nintendo Entertainment System) emulator written in C++ using SFML, supporting games that use no mapper or mappers 1, 2… | 43 | 5107 | abandoned |
| qTox/qTox qTox is a cross-platform instant messaging desktop client supporting encrypted text chat, voice and video calls, and file transfer over the… | 10 | 4978 | abandoned |
| chentao0707/SimplifyReader SimplifyReader is an Android client app built with Google Material Design, featuring five modules: news reading, image browsing, video watc… | 32 | 4544 | abandoned |
| accord-net/framework Accord.NET is a C# framework for .NET providing machine learning, statistics, computer vision, image and audio processing, and general scie… | 10 | 4535 | abandoned |
| xhzengAIB/MessageDisplayKit An Objective-C iOS library providing a WeChat-like instant messaging app experience, with UI components for sending text, pictures, audio, … | 23 | 4218 | abandoned |
| agermanidis/autosub Autosub is a Python command-line utility that auto-generates subtitles for video or audio files. It performs voice activity detection, tran… | 32 | 4191 | abandoned |
| SeriousCache/UABE UABE (Asset Bundle Extractor) is a Windows desktop editor for Unity .assets and AssetBundle files, supporting export/import of textures, te… | 10 | 4143 | abandoned |
| buriburisuri/speech-to-text-wavenet A TensorFlow implementation of end-to-end English speech recognition based on DeepMind's WaveNet architecture, trained with CTC loss on sen… | 32 | 4002 | abandoned |
| blackjack4494/yt-dlc youtube-dlc is a Python-based command-line media downloader and library for downloading videos from YouTube and other video platforms, fork… | 23 | 3041 | abandoned |
| zzw922cn/Automatic_Speech_Recognition An end-to-end automatic speech recognition system implemented in TensorFlow, supporting Mandarin and English with models like DeepSpeech2, … | 32 | 2831 | abandoned |
| replit/kaboom Kaboom is a JavaScript game library for building 2D games quickly in the browser, using a component-based system to compose game objects wi… | 10 | 2734 | abandoned |
| idootop/open-xiaoai Open-XiaoAI is a Rust-based client/server project that takes over the audio input and output of Xiaomi XiaoAI smart speakers (LX06 and OH2P… | 10 | 2595 | abandoned |
| nacker/LZEasemob3 An open-source iOS application that closely clones WeChat's UI and features (chat, contacts, Moments feed, voice messages, red packets) usi… | 32 | 1927 | abandoned |
| justLV/onju-voice A hackable AI home assistant platform that replaces the internals of a Google Nest Mini with a custom ESP32-S3 PCB, paired with a server th… | 63 | 1711 | abandoned |
| elevenlabs/elevenlabs-mcp The official ElevenLabs Model Context Protocol (MCP) server, exposing ElevenLabs text-to-speech, speech-to-text, voice design, and conversa… | 10 | 1532 | abandoned |
| artyshko/smd A Python application that downloads high-quality music from Spotify playlists, albums, and songs by fetching audio from sources like YouTub… | 23 | 1486 | abandoned |
| aFarkas/webshim Webshims Lib is a modular, capability-based polyfill-loading library that extends jQuery with HTML5 features in legacy browsers. It conditi… | 23 | 1399 | abandoned |
| Raival-e/Prism-File-Explorer Prism File Explorer is a modern, lightweight file manager for Android built entirely with Kotlin and Jetpack Compose, featuring a Material … | 64 | 1353 | abandoned |
| MediaCrush/MediaCrush MediaCrush is a self-hostable website for uploading and sharing images, audio, and video via fast, shareable links. It is a full web applic… | 32 | 1332 | abandoned |
← prev page 17 / 17