function: video-processing
1714 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| X-PLUG/mPLUG-Owl mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and… | 36 | 2539 | active |
| microsoft/ResearchStudio ResearchStudio is a Microsoft collection of AI agent skills that cover the entire research lifecycle, from an under-specified research dire… | 59 | 2523 | active |
| dino/dino Dino is a modern open-source XMPP (Jabber) chat client for Linux desktops, built with GTK4 and Vala. It supports end-to-end encryption via … | 76 | 2480 | active |
| synctv-org/synctv SyncTV is a self-hosted Rust server for real-time synchronized video watching with rooms, chat, and livestreaming (RTMP/HLS/HTTP-FLV/RTSP).… | 94 | 2474 | active |
| SickChill/sickchill SickChill is a self-hosted, cross-platform automatic video library manager (PVR) for TV shows written in Python with a web interface. It wa… | 67 | 2442 | active |
| volcengine/ai-app-lab AI App Lab from Volcano Engine (Volcengine) provides Arkitect, a high-code Python SDK for building LLM applications, plus Demohouse, a coll… | 83 | 2433 | active |
| roboflow/inference Roboflow Inference is a Python library and self-hostable inference server for deploying computer vision models on any computer or edge devi… | 91 | 2427 | active |
| snapotter-hq/SnapOtter SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities… | 80 | 2362 | active |
| facebookresearch/perception_models Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan… | 54 | 2353 | active |
| Xilinx/PYNQ PYNQ is an open-source Python framework from AMD/Xilinx for designing embedded systems on Zynq and other adaptive computing platforms (FPGA… | 78 | 2339 | active |
| stephengpope/no-code-architects-toolkit A self-hostable Flask-based API that consolidates common media processing tasks—video editing, captioning, audio conversion, transcription,… | 50 | 2339 | active |
| jeremyckahn/chitchatter Chitchatter is a free, open-source communication tool offering secure peer-to-peer chat, video, audio, and file sharing directly between br… | 75 | 2321 | active |
| samuelgursky/davinci-resolve-mcp A Model Context Protocol (MCP) server that lets AI assistants like Claude control DaVinci Resolve Studio through its official Scripting API… | 83 | 2300 | active |
| kerwincui/FastBee FastBee is a lightweight, full-stack open-source IoT platform built on Spring Boot with a built-in Netty MQTT broker, device management, th… | 66 | 2275 | active |
| OlafenwaMoses/ImageAI ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob… | 23 | 8877 | maintenance |
| hkchengrex/MMAudio MMAudio is a PyTorch-based model for generating synchronized audio from video and/or text inputs, using multimodal joint training across au… | 43 | 2264 | active |
| nv-tlabs/lyra Project Lyra is NVIDIA's open series of generative 3D world models, including Lyra 1.0 for feed-forward 3D/4D scene generation from a singl… | 59 | 2262 | active |
| rushindrasinha/youtube-shorts-pipeline Verticals v3 (repo youtube-shorts-pipeline) is a Python CLI that automates producing and publishing YouTube Shorts: it researches a topic, … | 61 | 2251 | active |
| Alpha-VLLM/Lumina-T2X Lumina-T2X is a unified framework for text-to-any-modality generation built on flow-based large diffusion transformers. It supports generat… | 28 | 2250 | active |
| andreknieriem/open-headunit Open Headunit is an Android app that turns an Android tablet, phone, or aftermarket head unit into an Android Auto receiver. It is a revive… | 79 | 2206 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2203 | active |
| baresip/baresip Baresip is a portable and modular SIP User-Agent written in C with audio and video call support. It provides a rich feature set including m… | 95 | 2202 | active |
| liangdabiao/Seedance2-Storyboard-Generator A Claude Code skill and workflow toolkit that converts novels or stories into multi-episode video series by generating four-act screenplays… | 55 | 2200 | active |
| TransWithAI/Faster-Whisper-TransWithAI-ChickenRice A high-performance audio/video transcription and translation application built on Faster Whisper, optimized for Japanese-to-Chinese transla… | 77 | 2196 | active |
| OpenVidu/openvidu OpenVidu is a self-hosted, open-source video conferencing platform and WebRTC SDK suite, built as a LiveKit fork with mediasoup as its SFU … | 91 | 2126 | active |
| liballeg/allegro5 Allegro 5 is a cross-platform C library for video game and multimedia programming, handling windows, input, graphics, audio, fonts, and vid… | 84 | 2120 | stable |
| ozgrozer/ai-renamer A Node.js CLI tool that uses local AI models (via Ollama or LM Studio) or OpenAI to intelligently rename files based on their contents, inc… | 26 | 2111 | active |
| MiniMax-AI/cli The official CLI for the MiniMax AI Platform, written in TypeScript, that generates text, images, video, speech, and music from the termina… | 81 | 2073 | active |
| ossia/score ossia score is a free, open-source interactive sequencer for audio-visual artists, designed to create interactive shows, installations, and… | 92 | 2053 | active |
| oil-oil/oil-motion Oil Motion is an agent-agnostic Skill that designs, generates, and integrates interactive web animations from AI-generated video. It handle… | 57 | 2041 | active |
| stackia/rtp2httpd A lightweight IPTV streaming relay server written in C that converts multicast RTP/UDP, RTSP, and HLS sources into unicast HTTP streams. It… | 90 | 2013 | active |
| ppwwyyxx/wechat-dump A Python-based tool that extracts and parses WeChat message history from a rooted Android phone, decoding the local message database and me… | 54 | 2010 | active |
| akiver/cs-demo-manager A free and open-source companion application for managing and analyzing Counter-Strike (CS2/CSGO) demo files. It extracts statistics, gener… | 97 | 1988 | active |
| open-mmlab/mmagic MMagic is OpenMMLab's toolbox for generative and multimodal AI image/video creation, built on PyTorch. It provides a large model zoo coveri… | 23 | 7457 | maintenance |
| sipsorcery-org/sipsorcery SIPSorcery is a C#/.NET library implementing SIP, WebRTC, RTP, ICE, STUN and SDP for building real-time communication applications. It is p… | 77 | 1930 | active |
| TypesettingTools/Aegisub Aegisub is a free, cross-platform advanced subtitle editor for creating and modifying subtitles. It provides audio-based timing, powerful s… | 79 | 1928 | active |
| lanyeeee/bilibili-video-downloader A cross-platform GUI desktop application built with Tauri for downloading videos, audio, subtitles, danmaku, and covers from Bilibili. It s… | 75 | 1912 | active |
| diffgram/diffgram Diffgram is a self-hosted AI datastore for managing schemas, BLOBs, and predictions, with built-in human supervision (data labeling), data … | 62 | 1909 | active |
| TheSpaghettiDetective/obico-server Obico Server is the self-hostable backend of the Obico smart 3D printing platform, providing AI-based print failure detection, webcam strea… | 76 | 1899 | active |
| szczyglis-dev/py-gpt PyGPT is an open-source, all-in-one desktop AI assistant for Linux, Windows, and Mac, written in Python. It supports chat, agents, vision, … | 92 | 1892 | active |
| lanbinleo/bili2text bili2text is a Python command-line tool that converts Bilibili videos into text transcripts from a link or BV number, handling download, au… | 56 | 1873 | active |
| MixLabPro/comfyui-mixlab-nodes A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec… | 67 | 1863 | active |
| thygate/stable-diffusion-webui-depthmap-script An extension for AUTOMATIC1111's Stable Diffusion WebUI that generates high-resolution depth maps from images using models like Marigold, M… | 32 | 1853 | active |
| RanFeng/clipsketch-ai ClipSketch AI is a web-based AI content creation workbench that imports videos from Bilibili and Xiaohongshu links, lets users frame-accura… | 43 | 1840 | active |
| boundless-large-model/boundless-world-model Boundless-World-Model (BWM) is a physically consistent, action-conditioned video world model built on Wan2.2-TI2V-5B that acts as a low-cos… | 59 | 1824 | active |
| martinlaxenaire/curtainsjs curtains.js is a lightweight vanilla WebGL JavaScript library that converts HTML DOM elements containing images, videos, and canvases into … | 39 | 1824 | active |
| whotto/Video_note_generator A Python tool that converts video URLs into polished Xiaohongshu (Little Red Book) notes and blog articles. It downloads videos, transcribe… | 43 | 1808 | active |
| Emu Series Emu3 is a suite of state-of-the-art multimodal models from BAAI trained solely with next-token prediction, tokenizing images, text, and vid… | 57 | 1778 | active |
| githubXiaowangzi/NP-Manager NP-Manager is an Android application for APK, DEX, JAR, Smali, PDF, and media file manipulation. It provides reverse-engineering features s… | 76 | 1758 | active |
| ThioJoe/Auto-Synced-Translated-Dubs A Python CLI tool that automatically translates video subtitles into multiple languages and generates AI voice dubbed audio tracks synced t… | 59 | 1747 | active |
| NVIDIA-NeMo/Curator NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for … | 86 | 1736 | active |
| skalesapp/skales Skales is a personal AI agent desktop and mobile application that runs locally on Windows, macOS, Linux, Android, and iOS, executing multi-… | 82 | 1727 | active |
| jau123/MeiGen-AI-Design-MCP An open-source MCP server that adds AI image and video generation capabilities to AI coding tools like Claude Code, Cursor, and Codex. It s… | 80 | 1719 | active |
| Novage/p2p-media-loader P2P Media Loader is an open-source JavaScript/TypeScript library that enables peer-to-peer delivery of live and on-demand HLS and MPEG-DASH… | 98 | 1712 | active |
| tddworks/baguette Baguette is a Swift CLI and WebSocket server that provides headless control of iOS simulators without opening Xcode or Simulator.app. It su… | 81 | 1704 | active |
| bytedance/Sa2VA Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understa… | 70 | 1666 | active |
| XueZeyue/DanceGRPO Official implementation of DanceGRPO, a framework applying Group Relative Policy Optimization (GRPO) to fine-tune visual generation models … | 40 | 1648 | active |
| JIA-Lab-research/ControlNeXt ControlNeXt is the official implementation of a controllable generation method for images and videos, built on Stable Diffusion XL, Stable … | 24 | 1646 | active |
| Benexl/yt-x A POSIX-compliant shell script that lets you browse YouTube and other yt-dlp-supported sites from the terminal using fzf or from an app lau… | 85 | 1645 | active |
| InterDigitalInc/CompressAI CompressAI is a PyTorch library and evaluation platform for end-to-end learned data compression research. It provides custom layers, entrop… | 73 | 1627 | active |
| Lynpoint/CyberVerse CyberVerse is an open-source, self-hosted framework for building real-time, voice-first AI agents with optional digital-human video (talkin… | 63 | 1612 | active |
| ml4a/ml4a ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li… | 32 | 1602 | active |
| X-LANCE/AniTalker AniTalker is the official PyTorch implementation of an ACM MM 2024 paper that animates a single static portrait into a vivid talking-face v… | 24 | 1598 | active |
| Drexubery/ViewCrafter ViewCrafter is a research codebase that uses video diffusion models to synthesize high-fidelity novel views of scenes from a single or spar… | 49 | 1587 | active |
| finnvoor/yap yap is a Swift CLI for on-device speech transcription of audio and video files using Apple's Speech.framework on macOS 26. It also supports… | 81 | 1586 | active |
| calzoneman/sync CyTube is a Node.js server and JavaScript/HTML web client for synchronizing online media playback across viewers in shared channels. Each c… | 51 | 1581 | active |
| Tencent/DepthCrafter DepthCrafter is a diffusion-based video depth estimation model from Tencent AI Lab that generates temporally consistent long depth sequence… | 38 | 1574 | active |
| MiniMax-AI/MiniMax-MCP The official MiniMax Model Context Protocol (MCP) server, written in Python, exposing MiniMax's text-to-speech, image generation, and video… | 64 | 1569 | active |
| jpush/aurora-imui Aurora IMUI is a general-purpose instant messaging UI component library providing MessageList and InputView components, independent of any … | 23 | 5696 | maintenance |
| shrimbly/node-banana Node Banana is an open-source, node-based visual workflow editor for building AI media generation pipelines. Users connect nodes on an infi… | 81 | 1550 | active |
| sepfy/libpeer libpeer is a portable WebRTC implementation written in C using BSD sockets, designed for IoT and embedded devices such as ESP32 and Raspber… | 77 | 1542 | active |
| roncoo/roncoo-education Roncoo Education (领课教育系统) is an open-source online education platform built with a Spring Cloud Alibaba microservices backend and Vue 3/Nux… | 68 | 1539 | active |
| tin2tin/Pallaidium Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE… | 75 | 1520 | active |
| microsoft/Mage Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage… | 57 | 1516 | active |
| NVlabs/describe-anything Describe Anything Model (DAM) is a vision-language model that generates detailed descriptions of user-specified regions in images and video… | 32 | 1514 | active |
| bytedance/SALMONN SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video… | 73 | 1513 | active |
| Jamailar/Beav Beav (formerly RedBox) is a local-first AI content operations workbench for social media creators, combining a desktop app and a Chrome/Edg… | 81 | 1509 | active |
| Taiizor/Sucrose Sucrose is a free, open-source wallpaper engine for Windows that renders interactive live wallpapers from GIFs, videos, URLs, web pages, Yo… | 92 | 1503 | active |
| TIGER-AI-Lab/TheoremExplainAgent An agentic AI system that generates long-form (5+ minute) Manim animation videos explaining mathematical and STEM theorems using LLM agents… | 35 | 1499 | active |
| technomancer702/nodecast-tv nodecast-tv is a self-hosted web-based IPTV player that streams Live TV, Movies, and Series from Xtream Codes or M3U providers directly in … | 58 | 1480 | active |
| code100x/cms An open-source content management system (LMS) that powers app.100xdevs.com, an online learning platform for coding cohorts and bootcamps. … | 68 | 1461 | active |
| p2r3/beheader A command-line tool that generates polyglot files - single files that are simultaneously valid images, videos, PDFs, ZIP archives, and HTML… | 48 | 1455 | active |
| ByteDance-Seed/m3-agent M3-Agent is a multimodal agent framework from ByteDance Seed that processes real-time visual and auditory inputs to build entity-centric lo… | 48 | 1445 | active |
| valentinfrlch/ha-llmvision LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live … | 90 | 1440 | active |
| Walter0807/MotionBERT Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske… | 65 | 1439 | active |
| krusemediallc/arcads-claude-code An agent skill pack and prompting library for generating AI marketing videos and images through the Arcads API, designed for use with Claud… | 55 | 1421 | active |
| jitsi/lib-jitsi-meet lib-jitsi-meet is a low-level JavaScript API library for building fully custom video conferencing experiences on top of Jitsi Meet infrastr… | 95 | 1419 | active |
| R3gm/SoniTranslate SoniTranslate is a Gradio-based web application that automatically dubs videos into other languages. It transcribes speech, translates it, … | 55 | 1410 | active |
| Arthi-chaud/Meelo Meelo is a self-hosted music streaming server designed for music collectors, similar to Plex or Jellyfin but focused on music. It offers ri… | 99 | 1398 | active |
| zackees/transcribe-anything A Python CLI app that transcribes local audio/video files or URLs (YouTube, Rumble, etc.) using multiple Whisper backends with automatic de… | 93 | 1393 | active |
| gerbera/gerbera Gerbera is a free, open-source UPnP/DLNA media server that streams digital media across a home network to compatible devices like TVs, game… | 86 | 1390 | active |
| zelon88/HRConvert2 HRConvert2 is a self-hosted, resource-aware file conversion server written in PHP that supports 488 file formats across documents, images, … | 99 | 1364 | active |
| ClimbSnail/HoloCubic_AIO HoloCubic_AIO is an open-source all-in-one third-party firmware for the HoloCubic ESP32-based holographic desktop device, bundling apps lik… | 29 | 1354 | active |
| yuanyuanxiang/SimpleRemoter SimpleRemoter (YAMA) is a C++ remote control suite derived from the Gh0st RAT codebase, providing remote desktop, file transfer, terminal, … | 60 | 1347 | active |
| wenqsun/DimensionX DimensionX is a research framework that generates photorealistic 3D and 4D scenes from a single image using controllable video diffusion mo… | 43 | 1333 | active |
| Nativ Nativ is a free, MIT-licensed macOS desktop application for running open AI models locally on Apple Silicon Macs, built on MLX-VLM. It prov… | 80 | 1330 | active |
| bytedance/Lance Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained … | 55 | 1329 | active |
| LLaVA-VL/LLaVA-NeXT LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision … | 64 | 4716 | maintenance |
| copperspice/copperspice CopperSpice is a set of cross-platform C++ libraries (Core, Gui, Network, Multimedia, SQL, OpenGL, Vulkan, WebKit, XML, and more) derived f… | 74 | 1324 | active |
| yukkcat/gemini-business2api A self-hosted gateway service that exposes Gemini Business through an OpenAI-compatible API, with multi-account load balancing and an admin… | 58 | 1321 | active |