function: video-processing
1714 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| VERT-sh/VERT VERT is an open-source web-based file converter that converts images, audio, documents, and video between 250+ formats. Non-video conversio… | 65 | 15402 | active |
| Chocobozzz/PeerTube PeerTube is a free, open-source, self-hosted video streaming platform that uses ActivityPub federation to connect independent video hosting… | 98 | 15277 | stable |
| kekingcn/kkFileView kkFileView is a self-hosted Spring Boot application that provides online preview of a very wide range of file formats, including Office doc… | 95 | 14595 | active |
| chatfire-AI/huobao-drama Huobao Drama is a self-hosted, AI-powered end-to-end short drama generation platform that automates the full pipeline from a one-sentence i… | 60 | 14141 | active |
| SubtitleEdit/subtitleedit Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac… | 94 | 13971 | active |
| LibreSpark/LibreTV LibreTV is a lightweight, free, self-hostable web application for searching and streaming video content aggregated from multiple third-part… | 10 | 13796 | active |
| DustinBrett/daedalOS daedalOS is a desktop environment that runs entirely in the browser, built with JavaScript. It provides a window manager, file system, task… | 77 | 13025 | active |
| modelscope/DiffSynth-Studio DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio … | 79 | 13003 | active |
| abus-aikorea/voice-pro Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, … | 84 | 12647 | active |
| NVIDIA/cosmos NVIDIA Cosmos is an open platform of omnimodal world foundation models, datasets, and tools for building Physical AI systems such as robots… | 72 | 11641 | active |
| facebookresearch/sam3 Official code for Meta's Segment Anything Model 3 (SAM 3), a unified foundation model for promptable segmentation in images and videos. It … | 63 | 11487 | active |
| owncast/owncast Owncast is a self-hosted, single-user live video streaming and chat server written in Go, offering an alternative to mainstream streaming p… | 88 | 11474 | active |
| TEN-framework/ten-framework TEN is an open-source framework for building real-time multimodal conversational AI agents, with a focus on low-latency voice assistants. I… | 85 | 11082 | active |
| voxel51/fiftyone FiftyOne is an open-source Python library and GUI app for building high-quality computer vision datasets and models. It enables visualizing… | 99 | 11042 | active |
| Baiyuetribe/paper2gui Paper2GUI (now branded 小白兔AI / Xiaobaitu AI) is a desktop AI toolbox application that packages 50+ AI research models into install-free GUI… | 23 | 10709 | active |
| iib0011/omni-tools OmniTools is a self-hosted web application bundling a large collection of browser-based utilities for images, video, audio, PDFs, text, dat… | 69 | 10088 | active |
| mrousavy/react-native-vision-camera A high-performance camera library for React Native offering photo/video capture, QR/barcode scanning, and JS worklet-based frame processors… | 99 | 9580 | active |
| LTX-2 Official Python package from Lightricks providing inference pipelines and LoRA training for LTX-2/LTX-2.5, an open-weights DiT-based founda… | 83 | 9260 | active |
| lipku/LiveTalking LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit… | 89 | 9238 | active |
| meetecho/janus-gateway Janus is an open-source, general-purpose WebRTC server written in C that sets up media communication with browsers and relays RTP/RTCP and … | 75 | 9156 | stable |
| NVlabs/Sana SANA is an efficiency-oriented PyTorch codebase for high-resolution text-to-image and text-to-video generation built on Linear Diffusion Tr… | 74 | 8833 | active |
| TeamWiseFlow/xiaobei Xiaobei is an open-source multi-agent system that automates social media marketing and customer acquisition for solo entrepreneurs and smal… | 90 | 8463 | active |
| XPixelGroup/BasicSR BasicSR is an open-source PyTorch toolbox for image and video restoration tasks such as super-resolution, denoising, deblurring, and JPEG a… | 23 | 8367 | stable |
| geekyutao/Inpaint-Anything Inpaint Anything combines Segment Anything (SAM) with inpainting models like LaMa and Stable Diffusion to remove, fill, or replace objects … | 65 | 7703 | active |
| chenyme/grok2api A self-hosted multi-account API gateway for Grok Build, Grok Web, and Grok Console, written in Go with a React frontend. It exposes Grok ca… | 84 | 7538 | active |
| LargeWorldModel/LWM Large World Model (LWM) is a family of open-source 7B-parameter multimodal autoregressive transformer models trained on long videos and boo… | 25 | 7425 | active |
| mediasoup mediasoup is a powerful WebRTC SFU (Selective Forwarding Unit) implemented as a Node.js module and Rust crate, with a C++ media worker subp… | 95 | 7345 | stable |
| yangchris11/samurai SAMURAI is the official implementation of a zero-shot visual object tracker built on top of Segment Anything Model 2 (SAM 2), using a motio… | 27 | 7112 | active |
| lightGallery lightGallery is a customizable, modular, dependency-free JavaScript lightbox gallery plugin for displaying images and videos on the web and… | 72 | 7049 | active |
| xushengfeng/eSearch eSearch is a cross-platform desktop application (Electron) combining screenshot capture, offline OCR based on PaddleOCR, screen search, tra… | 98 | 7036 | active |
| gaozhangmin/boxplayer BoxPlayer is a cross-platform desktop application that unifies multiple cloud drives (Aliyun Drive, Baidu, 115, Quark, OneDrive, etc.), loc… | 97 | 6878 | active |
| netless-io/flat Flat is the open-source Web, Windows, and macOS client of Agora Flat, an online classroom platform. It provides real-time interactive white… | 70 | 6415 | active |
| vllm-project/vllm-omni vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer… | 83 | 6369 | active |
| Forget-C/Jellyfish Jellyfish is a self-hosted, end-to-end production workspace for AI-generated short dramas, covering script input, storyboard breakdown, cha… | 73 | 6195 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| Tencent/libpag libpag is Tencent's official real-time rendering library for PAG (Portable Animated Graphics) files, a format exported from Adobe After Eff… | 98 | 5768 | active |
| DeepLabCut/DeepLabCut DeepLabCut is an open-source Python toolbox for markerless 2D and 3D pose estimation of user-defined body parts using deep neural networks … | 89 | 5745 | stable |
| Eventual-Inc/Daft Daft is a high-performance distributed data engine with a Python dataframe API, implemented in Rust, designed for AI and multimodal workloa… | 95 | 5730 | active |
| modstart-lib/aigcpanel AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, … | 88 | 5482 | active |
| Hillobar/Rope Rope is a GUI-focused desktop application for face swapping that implements the insightface inswapper_128 model. It offers batch swapping, … | 79 | 5367 | active |
| open-mmlab/mmaction2 MMAction2 is OpenMMLab's PyTorch-based toolbox and benchmark for video understanding, covering action recognition, temporal action localiza… | 55 | 5142 | active |
| signalwire/freeswitch FreeSWITCH is an open-source software-defined telecom stack that turns commodity servers into a full telephony platform for voice, video, a… | 94 | 5118 | stable |
| aiortc/aiortc aiortc is a Python library implementing WebRTC and ORTC on top of asyncio, with an API closely mirroring the JavaScript WebRTC API. It supp… | 74 | 5094 | active |
| facebookresearch/AugLy AugLy is a Python data augmentation library from Meta AI supporting audio, image, text, and video with over 100 augmentations. It focuses o… | 67 | 5089 | stable |
| mediacms-io/mediacms MediaCMS is an open source, self-hosted video and media CMS built with Django and React, exposing a REST API. It supports video, audio, ima… | 99 | 5083 | active |
| facebookresearch/co-tracker CoTracker is a transformer-based model from Meta AI and Oxford VGG that jointly tracks any point (pixel) across a video, handling occlusion… | 60 | 5080 | active |
| zju3dv/EasyMocap EasyMocap is an open-source Python toolbox for markerless human motion capture and novel view synthesis from RGB videos. It fits parametric… | 54 | 4783 | active |
| Steam-Headless/docker-steam-headless A headless Steam Docker image that runs a full Steam client and Xfce4 desktop on Linux with NVIDIA, AMD, or Intel GPU support. It enables r… | 58 | 4726 | active |
| miroslavpejic85/mirotalk MiroTalk P2P is a self-hosted, open-source WebRTC video conferencing platform that uses direct peer-to-peer connections for real-time video… | 77 | 4706 | active |
| AgnesAI-Labs/AgnesAI-Models Official gateway and model catalog for Agnes AI, providing OpenAI-compatible API access to multimodal foundation models covering text, imag… | 58 | 4705 | active |
| dankamongmen/notcurses Notcurses is a C library (with C++, Rust, and Python bindings) for building complex, vibrant textual user interfaces on modern terminal emu… | 76 | 4680 | active |
| LaoFeng-mouse/flyingmouse-format FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop… | 79 | 4601 | active |
| sensity-ai/dot dot (Deepfake Offensive Toolkit) is a Python tool that generates real-time, controllable deepfakes from a webcam feed and injects them into… | 23 | 4586 | active |
| royshil/obs-backgroundremoval An OBS Studio plugin that removes and replaces the background in portrait video using ONNX-based machine learning segmentation, acting as a… | 98 | 4492 | active |
| iperov/DeepFaceLab DeepFaceLab is the leading open-source Windows application for creating deepfakes, allowing users to swap, de-age, or replace faces and hea… | 10 | 19292 | maintenance |
| facebookresearch/jepa Official PyTorch implementation of V-JEPA, a self-supervised method for learning visual representations from video using a joint-embedding … | 29 | 4105 | active |
| hao-ai-lab/FastVideo FastVideo is a unified Python framework for post-training and real-time inference of video diffusion models, covering data preprocessing, f… | 80 | 4076 | active |
| QwenLM/Qwen2.5-Omni Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre… | 31 | 4074 | active |
| lllyasviel/Paints-UNDO Paints-UNDO is a family of deep learning models that take an image as input and generate the step-by-step drawing sequence (sketching, inki… | 40 | 4067 | active |
| umlx5h/LLPlayer LLPlayer is a Windows media player built for language learning, featuring dual subtitles, AI-generated subtitles via Whisper ASR, real-time… | 78 | 4049 | active |
| QwenLM/Qwen3-Omni Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an… | 52 | 3980 | active |
| Qsnh/meedu MeEdu is an open-source online school (knowledge commerce) system built with PHP/Laravel, MySQL, and Redis, supporting paid video courses, … | 90 | 3855 | active |
| google-research/scenic Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr… | 76 | 3821 | active |
| Perfare/AssetStudio AssetStudio is a Windows desktop GUI tool for exploring, extracting, and exporting assets and asset bundles from Unity games and projects. … | 10 | 15604 | maintenance |
| NExT-GPT/NExT-GPT NExT-GPT is an end-to-end any-to-any multimodal large language model that accepts and generates arbitrary combinations of text, image, vide… | 37 | 3638 | active |
| Haivision/srt SRT (Secure Reliable Transport) is an open-source transport protocol and C++ library for ultra-low-latency live video and audio streaming o… | 95 | 3585 | stable |
| MooreThreads/Moore-AnimateAnyone An open-source reproduction of AnimateAnyone that animates a character from a single reference image using pose sequences from a driving vi… | 26 | 3514 | active |
| PKU-YuanGroup/Video-LLaVA Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into… | 27 | 3500 | active |
| emacs-eaf/emacs-application-framework Emacs Application Framework (EAF) is an extensible framework that brings modern graphical capabilities to Emacs by bridging Emacs Lisp with… | 77 | 3487 | active |
| Kedreamix/Linly-Talker Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, … | 48 | 3436 | active |
| roflcoopter/viseron Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti… | 98 | 3431 | active |
| XiaoMi/xiaomi-miloco Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an… | 82 | 3292 | active |
| gozfree/gear-lib Gear-Lib is a collection of POSIX C libraries for IoT, embedded, and network service development, covering data structures, networking prot… | 23 | 3223 | active |
| wendy7756/AI-Video-Transcriber An open-source AI tool that transcribes, summarizes, and archives videos and podcasts from 30+ platforms (YouTube, TikTok, Bilibili, etc.) … | 62 | 3215 | active |
| xianfei/SysMocap SysMocap is a cross-platform, video-driven real-time motion capture system that animates 3D virtual characters from webcam footage. It rend… | 78 | 3199 | active |
| facebookresearch/tribev2 TRIBE v2 is a multimodal deep learning model from Meta AI that predicts fMRI brain responses to naturalistic video, audio, and text stimuli… | 55 | 3172 | active |
| miroslavpejic85/mirotalksfu MiroTalk SFU is a self-hosted, open-source WebRTC video conferencing platform built on the Mediasoup SFU architecture, positioning itself a… | 77 | 3087 | active |
| yuyuyzl/EasyVtuber EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int… | 62 | 3051 | active |
| TanStack/ai TanStack AI is a type-safe, provider-agnostic TypeScript SDK for building AI applications with streaming chat, tool calling, agents, struct… | 80 | 3028 | active |
| q191201771/lal LAL is an audio/video live streaming broadcast server written in Go, comparable to nginx-rtmp-module but with more features. It supports RT… | 23 | 3026 | active |
| SharpAI/DeepCamera DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r… | 86 | 3019 | active |
| sherlockchou86/VideoPipe VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates … | 54 | 2931 | active |
| InternLM/InternLM-XComposer InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u… | 38 | 2925 | active |
| datascale-ai/opentalking OpenTalking is an open-source Python framework for building real-time AI digital-human (talking avatar) conversation products. It orchestra… | 65 | 2897 | active |
| deforum/sd-webui-deforum Deforum is the official extension for AUTOMATIC1111's Stable Diffusion webui that generates AI animations from text prompts using keyframed… | 23 | 2859 | active |
| katelya77/KatelyaTV KatelyaTV is a self-hosted, cross-platform video aggregation player built on Next.js 14 and TypeScript, forked from MoonTV/LunaTV. It aggre… | 41 | 2857 | active |
| fqscfqj/Y2A-Auto Y2A-Auto is a self-hosted Python application that automates re-uploading YouTube videos to AcFun and bilibili. It handles the full pipeline… | 82 | 2845 | active |
| TheSmallHanCat/flow2api Flow2API is a self-hosted Python/FastAPI service that exposes an OpenAI- and Gemini-compatible API on top of Google Flow (VideoFX/ImageFX) … | 60 | 2834 | active |
| QwenLM/Qwen-MM-Plugins A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.… | 57 | 2777 | active |
| martinrotter/rssguard RSS Guard is a cross-platform desktop feed reader supporting RSS, ATOM, JSON and many web-based feed services, with a built-in podcast/medi… | 99 | 2731 | active |
| xdit-project/xDiT xDiT is a scalable inference engine for Diffusion Transformers (DiTs) that enables parallel deployment across multiple GPUs and machines. I… | 77 | 2699 | active |
| darkzOGx/youtube-automation-agent AgentTube is a self-hosted Node.js application that uses AI agents to run a YouTube channel end to end: researching topics, writing scripts… | 82 | 2666 | active |
| om-ai-lab/OmAgent OmAgent is a Python library for building multimodal language agents, wrapping worker orchestration, task queues, and graph-based workflow o… | 31 | 2665 | active |
| anliyuan/Ultralight-Digital-Human An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It… | 64 | 2627 | active |
| tl-open-source/tl-rtc-file A self-hostable web application that uses WebRTC peer-to-peer connections to transfer files, video, screen shares, live streams, and text b… | 23 | 2624 | active |
| morethanwords/tweb Telegram Web K is the open-source TypeScript web client powering web.telegram.org/k/, based on the original Webogram and actively patched a… | 77 | 2623 | active |
| kairos-agi/kairos Kairos is the official open-source implementation of a 4B-parameter native cross-embodiment world model that unifies video understanding, f… | 57 | 2598 | active |
| dreamzero0/dreamzero DreamZero is NVIDIA's World Action Model (WAM) that jointly predicts future video and actions from a pretrained video diffusion backbone, e… | 50 | 2593 | active |
| Tencent-Hunyuan/HY-World-2.0 HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from … | 58 | 2571 | active |
| fogsightai/fogsight Fogsight is an open-source AI agent and animation engine powered by large language models that turns abstract concepts or words into polish… | 51 | 2550 | active |