Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: video-processing

1714 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
DAMO-NLP-SG/VideoLLaMA2
VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro…
251307active
StarlightSearch/EmbedAnything
EmbedAnything is a high-performance, memory-safe embedding pipeline written in Rust (with Python bindings) that generates embeddings from t…
881305active
Henry-23/VideoChat
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late…
481303active
BelledonneCommunications/linphone-android
Linphone Android is an open-source SIP-based softphone application for voice and video calls over IP, plus instant messaging and conferenci…
981300active
Remocn/remocn
Remocn is a shadcn-style copy-paste component registry for Remotion, providing production-ready animations, transitions, kinetic typography…
581298active
kil0bit-kb/scrcpy-gui
ScrcpyGUI is a modern desktop GUI wrapper for scrcpy, built with Tauri v2, React 19, and Rust, for mirroring and controlling Android device…
821293active
buoyancy99/diffusion-forcing
Official research code for the NeurIPS paper 'Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion', implementing a metho…
651288active
stasel/WebRTC-iOS
A native iOS demo app written in Swift that shows the bare minimum needed to establish a peer-to-peer WebRTC connection, including audio an…
741278active
Renumics/spotlight
Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe…
931272active
fish2018/webhtv
WebHomeTV is an Android video streaming app (mobile and TV) forked from the FongMi/CatVod ecosystem, adding a customizable web-based homepa…
801264active
Fictionarry/ER-NeRF
ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr…
241260stable
roryclear/clearcam
Clearcam is a self-hosted Python NVR that adds AI object detection, tracking, mobile notifications, and semantic search to any RTSP securit…
861254active
zenstory-ai/drama-skills
A collection of ten agent skills for AI short-drama and comic-drama production, covering scripts, visual assets, storyboards, image/video p…
801240active
ryokun6/ryos
ryOS is a web-based desktop environment that recreates classic macOS and Windows interfaces in the browser, built with React and TypeScript…
841239active
zhistaredu/StarTraining
StarTraining (职星学院) is an open-source enterprise employee training and online education system built with Spring Boot and Vue, supporting o…
501232active
TheSmallHanCat/sora2api
A self-hosted OpenAI-compatible API gateway that wraps Sora's text-to-video and image generation capabilities behind standard /v1/chat/comp…
101232active
nicobailon/pi-web-access
A TypeScript extension for the Pi coding agent that adds web search, URL content extraction, GitHub repo cloning, PDF extraction, and YouTu…
831229active
suming77/SumTea_Android
SumTea is a WanAndroid client app for Android built with Kotlin, Jetpack, MVVM, coroutines, Flow, and Retrofit, featuring a componentized/m…
311228active
MotrixLab/SMPLer-X
Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an…
591220stable
thClaws/thClaws
thClaws is an open-source AI agent harness written in native Rust that ships as a single binary offering a desktop GUI, CLI, headless, and …
771208active
DachunKai/EvTexture
Official PyTorch implementation of EvTexture and EvTexture++, event-driven video super-resolution models that use event-camera signals to e…
541207active
AstraeLabs/VibraVid
VibraVid is a Python-based downloader for movies, series, anime, albums, and songs, supporting DASH, HLS, ISM, and MP4 streams including DR…
931202active
LifeArchiveProject/BilibiliHistoryFetcher
A Python/FastAPI backend tool that fetches, stores, and analyzes a user's Bilibili watch history, favorites, dynamics, comments, and intera…
811201active
EvolvingLMMs-Lab/LLaVA-OneVision-2
A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis…
721195active
reanimate/reanimate
Reanimate is a Haskell library for programmatically generating declarative 2D animations based on SVG graphics, inspired by 3b1b's manim. I…
341181active
DAMO-NLP-SG/VideoLLaMA3
VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de…
371179active
CH563/shot-easy-website
ShotEasy is a free online photo and screenshot toolkit built with Astro that runs entirely in the browser using WebAssembly. It offers scre…
701162active
MCG-NKU/E2FGVI
E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It…
321161stable
StreamerHelper/web-server
The backend service for StreamerHelper, a self-hosted livestream recording system. It polls live status from platforms like Bilibili, Huya,…
761153active
franklioxygen/MyTube
MyTube is a self-hosted web application that downloads videos from YouTube, Bilibili, Twitch, MissAV, and any yt-dlp-supported site, storin…
651153active
open-gigaai/giga-world-1
GigaWorld-1 is an open-source framework providing training, inference, data processing, checkpoint conversion, and LoRA merge workflows for…
541147active
PurpleDoubleD/locally-uncensored
Locally Uncensored is a free, open-source desktop AI studio (built with Tauri/TypeScript) that bundles uncensored local chat, a coding agen…
811146active
Kiteretsu77/APISR
APISR is a deep-learning based super-resolution tool that restores and enhances low-quality, low-resolution anime images and videos using t…
371135active
Materialious/Materialious
Materialious is a modern Material Design frontend client for YouTube and Invidious, available on Web, Desktop, Android, and Android TV. It …
881128active
MagicFoundation/Alcinoe
Alcinoe is a library of components and utilities for Delphi/FireMonkey (FMX) that helps developers build fast, modern, cross-platform appli…
911125active
HITsz-TMG/Uni-MoE
Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,…
681116active
ATH-MaaS/Pixelle-MCP
Pixelle MCP is an open-source omnimodal AIGC framework that converts ComfyUI workflows (local or RunningHub cloud) into MCP tools with zero…
441104active
welltop-cn/ComfyUI-TeaCache
A ComfyUI plugin integrating TeaCache, a training-free caching method that accelerates diffusion model inference by exploiting output diffe…
351092active
yerfor/Real3DPortrait
Official PyTorch implementation of Real3D-Portrait, an ICLR 2024 Spotlight paper for one-shot realistic 3D talking portrait synthesis. It g…
261091active
rhymes-ai/Aria
Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda…
231087active
facebookresearch/hiera
Hiera is the official PyTorch implementation of a hierarchical vision transformer from Meta AI (ICML 2023 Oral). It achieves state-of-the-a…
201074active
facebookresearch/CutLER
CutLER is a research codebase from Meta FAIR for training object detection and instance segmentation models without human annotations, usin…
661072active
AILab-CVC/UniRepLKNet
UniRepLKNet is a large-kernel ConvNet architecture (CVPR 2024, TPAMI 2025) that provides universal perception across image, audio, video, p…
431072stable
hmjz100/123panYouthMember
A Tampermonkey/Greasemonkey userscript that simulates 123pan (123 云盘) cloud drive membership features in the browser, including downloads o…
461070active
DevLARLEY/WidevineProxy2
A browser extension (Chrome and Firefox, Manifest V3) that proxies Widevine and ClearKey EME challenges and license messages, modifying cha…
861058active
sb-ai-lab/EmotiEffLib
EmotiEffLib (formerly HSEmotion) is a lightweight library for facial emotion and engagement recognition in photos and videos, available in …
651057active
yoyo-nb/Thin-Plate-Spline-Motion-Model
The official PyTorch implementation of the CVPR 2022 paper 'Thin-Plate Spline Motion Model for Image Animation'. It animates a source image…
323604maintenance
Tencent-Hunyuan/HunyuanVideo-Foley
HunyuanVideo-Foley is a multimodal diffusion model from Tencent Hunyuan that generates high-fidelity Foley sound effects synchronized with …
371052active
agents-flex/agents-flex
Agents-Flex is a lightweight, modular Java framework for building AI applications and agents, positioned as a Java counterpart to Spring AI…
931046active
maxzhang666/OneKeyVip
A multi-function browser userscript (compatible with Tampermonkey and ScriptCat) that bundles VIP video/music parsing, Bilibili cover fetch…
761030active
vastxie/99AI
99AI is a commercially viable, self-hostable AI web platform built with Vue and Node.js that bundles AI chat, image/video/music generation,…
351029active
wujunwei928/parse-video
A Go library and CLI tool that parses short-video share links from 25+ Chinese platforms (Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo, e…
811018active
towhee-io/towhee
Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s…
233454maintenance
ambrosiogabe/MathAnimation
A C++/OpenGL desktop application for creating mathematically accurate animations with a real-time GUI and audio preview, aiming to match Ma…
321017active
EvolvingLMMs-Lab/Otter
Otter is a multi-modal vision-language model built on OpenFlamingo, instruction-tuned on the MIMIC-IT dataset with image and video understa…
213436maintenance
siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System
A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro…
391001active
anandpawara/Real_Time_Image_Animation
A real-time Python application that animates a still image (e.g., a portrait) using facial motion from a live camera or video file, built o…
323248maintenance
DAMO-NLP-SG/Video-LLaMA
Video-LLaMA is an instruction-tuned audio-visual language model that extends LLaMA with video and audio understanding via cross-modal pretr…
293139maintenance
lynckia/licode
Licode is an open-source WebRTC communications platform for hosting your own videoconference provider, built around the C++ Erizo MCU, a No…
673134maintenance
StarRTC
StarRTC is a free cross-platform real-time communication SDK and self-hostable server suite providing instant messaging, one-to-one video c…
233075maintenance
huangruiLearn/flutter_hrlweibo
A Weibo (Chinese microblog) client clone built with Flutter, replicating roughly 80% of Weibo's UI across dozens of screens. It includes ho…
322864maintenance
microsoft/NUWA
Microsoft's official research repository for the NUWA family of multimodal generative models, a unified 3D transformer pipeline for visual …
102791maintenance
zzh8829/yolov3-tf2
A clean implementation of YOLOv3 and YOLOv3-tiny object detection in TensorFlow 2.0, with pre-trained Darknet weight conversion, inference,…
322513maintenance
lbryio/lbry-android
The official LBRY Android app, a mobile browser and wallet for the LBRY decentralized content network. It lets users discover, view, publis…
232396maintenance
facebookresearch/frankmocap
FrankMocap is a single-view 3D motion capture system from Facebook AI Research that estimates 3D pose for body, hands, and whole body (body…
102294maintenance
qianqianwang68/omnimotion
OmniMotion is a PyTorch implementation of the ICCV 2023 paper 'Tracking Everything Everywhere All at Once', which tracks every point in a v…
292268maintenance
vercel/virtual-event-starter-kit
An open-source Next.js starter kit for hosting virtual events and conferences, used to run Next.js Conf 2020 with ~40,000 attendees. It pro…
102178maintenance
google-deepmind/kinetics-i3d
A repository of pre-trained Inflated 3D Convnet (I3D) models for video action classification, trained on the Kinetics dataset, released alo…
321838maintenance
openai/Video-Pre-Training
OpenAI's Video PreTraining (VPT) codebase for learning Minecraft agents by watching unlabeled online videos, including behavioral cloning a…
411737maintenance
NVIDIA/Cosmos-Tokenizer
NVIDIA Cosmos Tokenizer is a suite of neural tokenizers for images and videos that convert visual data into continuous latents or discrete …
101731maintenance
Lightning-Universe/lightning-flash
Lightning Flash is a high-level PyTorch library built on PyTorch Lightning that provides ready-made 'recipes' for over 15 AI tasks across 7…
101722maintenance
invictus717/MetaTransformer
Meta-Transformer is a research framework for unified multimodal learning that maps inputs from 12 modalities (text, images, point clouds, a…
191647maintenance
facebookresearch/consistent_depth
A research library from Facebook AI Research implementing Consistent Video Depth Estimation (SIGGRAPH 2020). It reconstructs dense, flicker…
101634maintenance
xinntao/EDVR
EDVR is the winning solution of the NTIRE19 video restoration challenges, built on enhanced deformable convolutional networks. The repo is …
321577maintenance
sniklaus/3d-ken-burns
A PyTorch reference implementation of the 3D Ken Burns Effect from a Single Image paper, which animates a still photo with a virtual camera…
701569maintenance
Javacr/PyQt5-YOLOv5
A desktop GUI application built with PyQt5 that wraps YOLOv5 (v6.1) object detection models. It supports running detection on images, video…
321547maintenance
JingyunLiang/VRT
VRT is the official PyTorch implementation of the paper 'VRT: A Video Restoration Transformer', a transformer-based model for video restora…
231546maintenance
VideoData/DY-Data
A collection of Douyin (Chinese TikTok) scraping tools and API source code covering search, user, video, live stream, comments, danmaku, an…
321510maintenance
mikaelzero/mojito
An Android library that provides WeChat/Bilibili-style image and video viewer transitions, including drag-to-dismiss, large image, long ima…
101502maintenance
mps-youtube/pafy
Pafy is a Python library for retrieving YouTube video metadata and downloading video or audio streams at requested resolutions and formats.…
231416maintenance
guangqiang-liu/OneM
OneM is a comprehensive React Native app combining magazine browsing, music playback, and video playback, built with Redux state management…
321375maintenance
msracver/Deep-Feature-Flow
Official MXNet implementation of Deep Feature Flow (CVPR 2017), an end-to-end framework for video recognition such as object detection and …
321315maintenance
vivianli-me/ReactNativeOne
A React Native clone of the Chinese literary lifestyle app 'ONE·一个', covering picture-text, reading, music, and movie sections with ~80% co…
701284maintenance
rotemtzaban/STIT
STIT (Stitch it in Time) is a research implementation of a GAN-based framework for semantic editing of faces in real videos, based on the p…
321197maintenance
andrewkirillov/AForge.NET
AForge.NET is an open-source C# framework for computer vision and artificial intelligence, comprising libraries such as AForge.Imaging, AFo…
321151maintenance
hotshotco/Hotshot-XL
Hotshot-XL is an AI text-to-GIF model built to work alongside Stable Diffusion XL, generating 1-second GIFs at 8 FPS. It supports any fine-…
271111maintenance
qiucheng025/zao-
A Python deep learning tool that identifies and swaps faces in images and videos, with extract, train, and convert workflows plus an option…
321104maintenance
xunlu129/teriteri-client
Teriteri is a Vue3-based web client for a Bilibili-style danmaku (bullet comment) video sharing platform, built as a graduation project wit…
291095maintenance
YudongGuo/AD-NeRF
A PyTorch implementation of AD-NeRF, an ICCV 2021 paper that synthesizes talking-head videos by driving neural radiance fields with audio i…
321072maintenance
microsoft/VideoX
VideoX is a collection of Microsoft's video cross-modal understanding models, including X-CLIP for video-language recognition, 2D-TAN and M…
321071maintenance
ArrowLuo/CLIP4Clip
Official PyTorch implementation of the CLIP4Clip paper, a video-text retrieval model that transfers CLIP knowledge to end-to-end video clip…
231031maintenance
Everlyn-Labs/Everlyn-1
Everlyn-1 is an open autoregressive foundational video AI model from Everlyn Labs, accompanied by research on video compression/tokenizatio…
222892experimental
aleksilassila/reiverr
Reiverr is a self-hosted web application providing a unified interface for discovering movies and TV shows via TMDB and streaming content f…
612341experimental
etched-ai/open-oasis
Inference code and model weights for Oasis 500M, an interactive world model from Decart and Etched that generates gameplay video autoregres…
222123experimental
lucidrains/make-a-video-pytorch
A PyTorch library implementing Make-A-Video, Meta AI's text-to-video generation approach, built around pseudo-3d (axial) convolutions and s…
231986experimental
lyuchenyang/Macaw-LLM
Macaw-LLM is a multi-modal language modeling framework that integrates image, video, audio, and text data, built on CLIP, Whisper, and LLaM…
291591experimental
orca-wm/Orca
Orca is a general world foundation model from BAAI centered on Next-State-Prediction, learning a unified world latent space from visual and…
571038experimental
HumanMLLM/R1-Omni
R1-Omni is a research project applying Reinforcement Learning with Verifiable Reward (RLVR) to an omni-multimodal large language model for …
261022experimental
qTox/qTox
qTox is a cross-platform instant messaging desktop client supporting encrypted text chat, voice and video calls, and file transfer over the…
104978abandoned
xhzengAIB/MessageDisplayKit
An Objective-C iOS library providing a WeChat-like instant messaging app experience, with UI components for sending text, pictures, audio, …
234218abandoned

← prev page 17 / 18 next →