function: video-processing
1714 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| DAMO-NLP-SG/VideoLLaMA2 VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro… | 25 | 1307 | active |
| StarlightSearch/EmbedAnything EmbedAnything is a high-performance, memory-safe embedding pipeline written in Rust (with Python bindings) that generates embeddings from t… | 88 | 1305 | active |
| Henry-23/VideoChat A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late… | 48 | 1303 | active |
| BelledonneCommunications/linphone-android Linphone Android is an open-source SIP-based softphone application for voice and video calls over IP, plus instant messaging and conferenci… | 98 | 1300 | active |
| Remocn/remocn Remocn is a shadcn-style copy-paste component registry for Remotion, providing production-ready animations, transitions, kinetic typography… | 58 | 1298 | active |
| kil0bit-kb/scrcpy-gui ScrcpyGUI is a modern desktop GUI wrapper for scrcpy, built with Tauri v2, React 19, and Rust, for mirroring and controlling Android device… | 82 | 1293 | active |
| buoyancy99/diffusion-forcing Official research code for the NeurIPS paper 'Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion', implementing a metho… | 65 | 1288 | active |
| stasel/WebRTC-iOS A native iOS demo app written in Swift that shows the bare minimum needed to establish a peer-to-peer WebRTC connection, including audio an… | 74 | 1278 | active |
| Renumics/spotlight Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe… | 93 | 1272 | active |
| fish2018/webhtv WebHomeTV is an Android video streaming app (mobile and TV) forked from the FongMi/CatVod ecosystem, adding a customizable web-based homepa… | 80 | 1264 | active |
| Fictionarry/ER-NeRF ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr… | 24 | 1260 | stable |
| roryclear/clearcam Clearcam is a self-hosted Python NVR that adds AI object detection, tracking, mobile notifications, and semantic search to any RTSP securit… | 86 | 1254 | active |
| zenstory-ai/drama-skills A collection of ten agent skills for AI short-drama and comic-drama production, covering scripts, visual assets, storyboards, image/video p… | 80 | 1240 | active |
| ryokun6/ryos ryOS is a web-based desktop environment that recreates classic macOS and Windows interfaces in the browser, built with React and TypeScript… | 84 | 1239 | active |
| zhistaredu/StarTraining StarTraining (职星学院) is an open-source enterprise employee training and online education system built with Spring Boot and Vue, supporting o… | 50 | 1232 | active |
| TheSmallHanCat/sora2api A self-hosted OpenAI-compatible API gateway that wraps Sora's text-to-video and image generation capabilities behind standard /v1/chat/comp… | 10 | 1232 | active |
| nicobailon/pi-web-access A TypeScript extension for the Pi coding agent that adds web search, URL content extraction, GitHub repo cloning, PDF extraction, and YouTu… | 83 | 1229 | active |
| suming77/SumTea_Android SumTea is a WanAndroid client app for Android built with Kotlin, Jetpack, MVVM, coroutines, Flow, and Retrofit, featuring a componentized/m… | 31 | 1228 | active |
| MotrixLab/SMPLer-X Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an… | 59 | 1220 | stable |
| thClaws/thClaws thClaws is an open-source AI agent harness written in native Rust that ships as a single binary offering a desktop GUI, CLI, headless, and … | 77 | 1208 | active |
| DachunKai/EvTexture Official PyTorch implementation of EvTexture and EvTexture++, event-driven video super-resolution models that use event-camera signals to e… | 54 | 1207 | active |
| AstraeLabs/VibraVid VibraVid is a Python-based downloader for movies, series, anime, albums, and songs, supporting DASH, HLS, ISM, and MP4 streams including DR… | 93 | 1202 | active |
| LifeArchiveProject/BilibiliHistoryFetcher A Python/FastAPI backend tool that fetches, stores, and analyzes a user's Bilibili watch history, favorites, dynamics, comments, and intera… | 81 | 1201 | active |
| EvolvingLMMs-Lab/LLaVA-OneVision-2 A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis… | 72 | 1195 | active |
| reanimate/reanimate Reanimate is a Haskell library for programmatically generating declarative 2D animations based on SVG graphics, inspired by 3b1b's manim. I… | 34 | 1181 | active |
| DAMO-NLP-SG/VideoLLaMA3 VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de… | 37 | 1179 | active |
| CH563/shot-easy-website ShotEasy is a free online photo and screenshot toolkit built with Astro that runs entirely in the browser using WebAssembly. It offers scre… | 70 | 1162 | active |
| MCG-NKU/E2FGVI E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It… | 32 | 1161 | stable |
| StreamerHelper/web-server The backend service for StreamerHelper, a self-hosted livestream recording system. It polls live status from platforms like Bilibili, Huya,… | 76 | 1153 | active |
| franklioxygen/MyTube MyTube is a self-hosted web application that downloads videos from YouTube, Bilibili, Twitch, MissAV, and any yt-dlp-supported site, storin… | 65 | 1153 | active |
| open-gigaai/giga-world-1 GigaWorld-1 is an open-source framework providing training, inference, data processing, checkpoint conversion, and LoRA merge workflows for… | 54 | 1147 | active |
| PurpleDoubleD/locally-uncensored Locally Uncensored is a free, open-source desktop AI studio (built with Tauri/TypeScript) that bundles uncensored local chat, a coding agen… | 81 | 1146 | active |
| Kiteretsu77/APISR APISR is a deep-learning based super-resolution tool that restores and enhances low-quality, low-resolution anime images and videos using t… | 37 | 1135 | active |
| Materialious/Materialious Materialious is a modern Material Design frontend client for YouTube and Invidious, available on Web, Desktop, Android, and Android TV. It … | 88 | 1128 | active |
| MagicFoundation/Alcinoe Alcinoe is a library of components and utilities for Delphi/FireMonkey (FMX) that helps developers build fast, modern, cross-platform appli… | 91 | 1125 | active |
| HITsz-TMG/Uni-MoE Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,… | 68 | 1116 | active |
| ATH-MaaS/Pixelle-MCP Pixelle MCP is an open-source omnimodal AIGC framework that converts ComfyUI workflows (local or RunningHub cloud) into MCP tools with zero… | 44 | 1104 | active |
| welltop-cn/ComfyUI-TeaCache A ComfyUI plugin integrating TeaCache, a training-free caching method that accelerates diffusion model inference by exploiting output diffe… | 35 | 1092 | active |
| yerfor/Real3DPortrait Official PyTorch implementation of Real3D-Portrait, an ICLR 2024 Spotlight paper for one-shot realistic 3D talking portrait synthesis. It g… | 26 | 1091 | active |
| rhymes-ai/Aria Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda… | 23 | 1087 | active |
| facebookresearch/hiera Hiera is the official PyTorch implementation of a hierarchical vision transformer from Meta AI (ICML 2023 Oral). It achieves state-of-the-a… | 20 | 1074 | active |
| facebookresearch/CutLER CutLER is a research codebase from Meta FAIR for training object detection and instance segmentation models without human annotations, usin… | 66 | 1072 | active |
| AILab-CVC/UniRepLKNet UniRepLKNet is a large-kernel ConvNet architecture (CVPR 2024, TPAMI 2025) that provides universal perception across image, audio, video, p… | 43 | 1072 | stable |
| hmjz100/123panYouthMember A Tampermonkey/Greasemonkey userscript that simulates 123pan (123 云盘) cloud drive membership features in the browser, including downloads o… | 46 | 1070 | active |
| DevLARLEY/WidevineProxy2 A browser extension (Chrome and Firefox, Manifest V3) that proxies Widevine and ClearKey EME challenges and license messages, modifying cha… | 86 | 1058 | active |
| sb-ai-lab/EmotiEffLib EmotiEffLib (formerly HSEmotion) is a lightweight library for facial emotion and engagement recognition in photos and videos, available in … | 65 | 1057 | active |
| yoyo-nb/Thin-Plate-Spline-Motion-Model The official PyTorch implementation of the CVPR 2022 paper 'Thin-Plate Spline Motion Model for Image Animation'. It animates a source image… | 32 | 3604 | maintenance |
| Tencent-Hunyuan/HunyuanVideo-Foley HunyuanVideo-Foley is a multimodal diffusion model from Tencent Hunyuan that generates high-fidelity Foley sound effects synchronized with … | 37 | 1052 | active |
| agents-flex/agents-flex Agents-Flex is a lightweight, modular Java framework for building AI applications and agents, positioned as a Java counterpart to Spring AI… | 93 | 1046 | active |
| maxzhang666/OneKeyVip A multi-function browser userscript (compatible with Tampermonkey and ScriptCat) that bundles VIP video/music parsing, Bilibili cover fetch… | 76 | 1030 | active |
| vastxie/99AI 99AI is a commercially viable, self-hostable AI web platform built with Vue and Node.js that bundles AI chat, image/video/music generation,… | 35 | 1029 | active |
| wujunwei928/parse-video A Go library and CLI tool that parses short-video share links from 25+ Chinese platforms (Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo, e… | 81 | 1018 | active |
| towhee-io/towhee Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s… | 23 | 3454 | maintenance |
| ambrosiogabe/MathAnimation A C++/OpenGL desktop application for creating mathematically accurate animations with a real-time GUI and audio preview, aiming to match Ma… | 32 | 1017 | active |
| EvolvingLMMs-Lab/Otter Otter is a multi-modal vision-language model built on OpenFlamingo, instruction-tuned on the MIMIC-IT dataset with image and video understa… | 21 | 3436 | maintenance |
| siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro… | 39 | 1001 | active |
| anandpawara/Real_Time_Image_Animation A real-time Python application that animates a still image (e.g., a portrait) using facial motion from a live camera or video file, built o… | 32 | 3248 | maintenance |
| DAMO-NLP-SG/Video-LLaMA Video-LLaMA is an instruction-tuned audio-visual language model that extends LLaMA with video and audio understanding via cross-modal pretr… | 29 | 3139 | maintenance |
| lynckia/licode Licode is an open-source WebRTC communications platform for hosting your own videoconference provider, built around the C++ Erizo MCU, a No… | 67 | 3134 | maintenance |
| StarRTC StarRTC is a free cross-platform real-time communication SDK and self-hostable server suite providing instant messaging, one-to-one video c… | 23 | 3075 | maintenance |
| huangruiLearn/flutter_hrlweibo A Weibo (Chinese microblog) client clone built with Flutter, replicating roughly 80% of Weibo's UI across dozens of screens. It includes ho… | 32 | 2864 | maintenance |
| microsoft/NUWA Microsoft's official research repository for the NUWA family of multimodal generative models, a unified 3D transformer pipeline for visual … | 10 | 2791 | maintenance |
| zzh8829/yolov3-tf2 A clean implementation of YOLOv3 and YOLOv3-tiny object detection in TensorFlow 2.0, with pre-trained Darknet weight conversion, inference,… | 32 | 2513 | maintenance |
| lbryio/lbry-android The official LBRY Android app, a mobile browser and wallet for the LBRY decentralized content network. It lets users discover, view, publis… | 23 | 2396 | maintenance |
| facebookresearch/frankmocap FrankMocap is a single-view 3D motion capture system from Facebook AI Research that estimates 3D pose for body, hands, and whole body (body… | 10 | 2294 | maintenance |
| qianqianwang68/omnimotion OmniMotion is a PyTorch implementation of the ICCV 2023 paper 'Tracking Everything Everywhere All at Once', which tracks every point in a v… | 29 | 2268 | maintenance |
| vercel/virtual-event-starter-kit An open-source Next.js starter kit for hosting virtual events and conferences, used to run Next.js Conf 2020 with ~40,000 attendees. It pro… | 10 | 2178 | maintenance |
| google-deepmind/kinetics-i3d A repository of pre-trained Inflated 3D Convnet (I3D) models for video action classification, trained on the Kinetics dataset, released alo… | 32 | 1838 | maintenance |
| openai/Video-Pre-Training OpenAI's Video PreTraining (VPT) codebase for learning Minecraft agents by watching unlabeled online videos, including behavioral cloning a… | 41 | 1737 | maintenance |
| NVIDIA/Cosmos-Tokenizer NVIDIA Cosmos Tokenizer is a suite of neural tokenizers for images and videos that convert visual data into continuous latents or discrete … | 10 | 1731 | maintenance |
| Lightning-Universe/lightning-flash Lightning Flash is a high-level PyTorch library built on PyTorch Lightning that provides ready-made 'recipes' for over 15 AI tasks across 7… | 10 | 1722 | maintenance |
| invictus717/MetaTransformer Meta-Transformer is a research framework for unified multimodal learning that maps inputs from 12 modalities (text, images, point clouds, a… | 19 | 1647 | maintenance |
| facebookresearch/consistent_depth A research library from Facebook AI Research implementing Consistent Video Depth Estimation (SIGGRAPH 2020). It reconstructs dense, flicker… | 10 | 1634 | maintenance |
| xinntao/EDVR EDVR is the winning solution of the NTIRE19 video restoration challenges, built on enhanced deformable convolutional networks. The repo is … | 32 | 1577 | maintenance |
| sniklaus/3d-ken-burns A PyTorch reference implementation of the 3D Ken Burns Effect from a Single Image paper, which animates a still photo with a virtual camera… | 70 | 1569 | maintenance |
| Javacr/PyQt5-YOLOv5 A desktop GUI application built with PyQt5 that wraps YOLOv5 (v6.1) object detection models. It supports running detection on images, video… | 32 | 1547 | maintenance |
| JingyunLiang/VRT VRT is the official PyTorch implementation of the paper 'VRT: A Video Restoration Transformer', a transformer-based model for video restora… | 23 | 1546 | maintenance |
| VideoData/DY-Data A collection of Douyin (Chinese TikTok) scraping tools and API source code covering search, user, video, live stream, comments, danmaku, an… | 32 | 1510 | maintenance |
| mikaelzero/mojito An Android library that provides WeChat/Bilibili-style image and video viewer transitions, including drag-to-dismiss, large image, long ima… | 10 | 1502 | maintenance |
| mps-youtube/pafy Pafy is a Python library for retrieving YouTube video metadata and downloading video or audio streams at requested resolutions and formats.… | 23 | 1416 | maintenance |
| guangqiang-liu/OneM OneM is a comprehensive React Native app combining magazine browsing, music playback, and video playback, built with Redux state management… | 32 | 1375 | maintenance |
| msracver/Deep-Feature-Flow Official MXNet implementation of Deep Feature Flow (CVPR 2017), an end-to-end framework for video recognition such as object detection and … | 32 | 1315 | maintenance |
| vivianli-me/ReactNativeOne A React Native clone of the Chinese literary lifestyle app 'ONE·一个', covering picture-text, reading, music, and movie sections with ~80% co… | 70 | 1284 | maintenance |
| rotemtzaban/STIT STIT (Stitch it in Time) is a research implementation of a GAN-based framework for semantic editing of faces in real videos, based on the p… | 32 | 1197 | maintenance |
| andrewkirillov/AForge.NET AForge.NET is an open-source C# framework for computer vision and artificial intelligence, comprising libraries such as AForge.Imaging, AFo… | 32 | 1151 | maintenance |
| hotshotco/Hotshot-XL Hotshot-XL is an AI text-to-GIF model built to work alongside Stable Diffusion XL, generating 1-second GIFs at 8 FPS. It supports any fine-… | 27 | 1111 | maintenance |
| qiucheng025/zao- A Python deep learning tool that identifies and swaps faces in images and videos, with extract, train, and convert workflows plus an option… | 32 | 1104 | maintenance |
| xunlu129/teriteri-client Teriteri is a Vue3-based web client for a Bilibili-style danmaku (bullet comment) video sharing platform, built as a graduation project wit… | 29 | 1095 | maintenance |
| YudongGuo/AD-NeRF A PyTorch implementation of AD-NeRF, an ICCV 2021 paper that synthesizes talking-head videos by driving neural radiance fields with audio i… | 32 | 1072 | maintenance |
| microsoft/VideoX VideoX is a collection of Microsoft's video cross-modal understanding models, including X-CLIP for video-language recognition, 2D-TAN and M… | 32 | 1071 | maintenance |
| ArrowLuo/CLIP4Clip Official PyTorch implementation of the CLIP4Clip paper, a video-text retrieval model that transfers CLIP knowledge to end-to-end video clip… | 23 | 1031 | maintenance |
| Everlyn-Labs/Everlyn-1 Everlyn-1 is an open autoregressive foundational video AI model from Everlyn Labs, accompanied by research on video compression/tokenizatio… | 22 | 2892 | experimental |
| aleksilassila/reiverr Reiverr is a self-hosted web application providing a unified interface for discovering movies and TV shows via TMDB and streaming content f… | 61 | 2341 | experimental |
| etched-ai/open-oasis Inference code and model weights for Oasis 500M, an interactive world model from Decart and Etched that generates gameplay video autoregres… | 22 | 2123 | experimental |
| lucidrains/make-a-video-pytorch A PyTorch library implementing Make-A-Video, Meta AI's text-to-video generation approach, built around pseudo-3d (axial) convolutions and s… | 23 | 1986 | experimental |
| lyuchenyang/Macaw-LLM Macaw-LLM is a multi-modal language modeling framework that integrates image, video, audio, and text data, built on CLIP, Whisper, and LLaM… | 29 | 1591 | experimental |
| orca-wm/Orca Orca is a general world foundation model from BAAI centered on Next-State-Prediction, learning a unified world latent space from visual and… | 57 | 1038 | experimental |
| HumanMLLM/R1-Omni R1-Omni is a research project applying Reinforcement Learning with Verifiable Reward (RLVR) to an omni-multimodal large language model for … | 26 | 1022 | experimental |
| qTox/qTox qTox is a cross-platform instant messaging desktop client supporting encrypted text chat, voice and video calls, and file transfer over the… | 10 | 4978 | abandoned |
| xhzengAIB/MessageDisplayKit An Objective-C iOS library providing a WeChat-like instant messaging app experience, with UI components for sending text, pictures, audio, … | 23 | 4218 | abandoned |