Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: video-processing

1714 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
VERT-sh/VERT
VERT is an open-source web-based file converter that converts images, audio, documents, and video between 250+ formats. Non-video conversio…
6515402active
Chocobozzz/PeerTube
PeerTube is a free, open-source, self-hosted video streaming platform that uses ActivityPub federation to connect independent video hosting…
9815277stable
kekingcn/kkFileView
kkFileView is a self-hosted Spring Boot application that provides online preview of a very wide range of file formats, including Office doc…
9514595active
chatfire-AI/huobao-drama
Huobao Drama is a self-hosted, AI-powered end-to-end short drama generation platform that automates the full pipeline from a one-sentence i…
6014141active
SubtitleEdit/subtitleedit
Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac…
9413971active
LibreSpark/LibreTV
LibreTV is a lightweight, free, self-hostable web application for searching and streaming video content aggregated from multiple third-part…
1013796active
DustinBrett/daedalOS
daedalOS is a desktop environment that runs entirely in the browser, built with JavaScript. It provides a window manager, file system, task…
7713025active
modelscope/DiffSynth-Studio
DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio …
7913003active
abus-aikorea/voice-pro
Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, …
8412647active
NVIDIA/cosmos
NVIDIA Cosmos is an open platform of omnimodal world foundation models, datasets, and tools for building Physical AI systems such as robots…
7211641active
facebookresearch/sam3
Official code for Meta's Segment Anything Model 3 (SAM 3), a unified foundation model for promptable segmentation in images and videos. It …
6311487active
owncast/owncast
Owncast is a self-hosted, single-user live video streaming and chat server written in Go, offering an alternative to mainstream streaming p…
8811474active
TEN-framework/ten-framework
TEN is an open-source framework for building real-time multimodal conversational AI agents, with a focus on low-latency voice assistants. I…
8511082active
voxel51/fiftyone
FiftyOne is an open-source Python library and GUI app for building high-quality computer vision datasets and models. It enables visualizing…
9911042active
Baiyuetribe/paper2gui
Paper2GUI (now branded 小白兔AI / Xiaobaitu AI) is a desktop AI toolbox application that packages 50+ AI research models into install-free GUI…
2310709active
iib0011/omni-tools
OmniTools is a self-hosted web application bundling a large collection of browser-based utilities for images, video, audio, PDFs, text, dat…
6910088active
mrousavy/react-native-vision-camera
A high-performance camera library for React Native offering photo/video capture, QR/barcode scanning, and JS worklet-based frame processors…
999580active
LTX-2
Official Python package from Lightricks providing inference pipelines and LoRA training for LTX-2/LTX-2.5, an open-weights DiT-based founda…
839260active
lipku/LiveTalking
LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit…
899238active
meetecho/janus-gateway
Janus is an open-source, general-purpose WebRTC server written in C that sets up media communication with browsers and relays RTP/RTCP and …
759156stable
NVlabs/Sana
SANA is an efficiency-oriented PyTorch codebase for high-resolution text-to-image and text-to-video generation built on Linear Diffusion Tr…
748833active
TeamWiseFlow/xiaobei
Xiaobei is an open-source multi-agent system that automates social media marketing and customer acquisition for solo entrepreneurs and smal…
908463active
XPixelGroup/BasicSR
BasicSR is an open-source PyTorch toolbox for image and video restoration tasks such as super-resolution, denoising, deblurring, and JPEG a…
238367stable
geekyutao/Inpaint-Anything
Inpaint Anything combines Segment Anything (SAM) with inpainting models like LaMa and Stable Diffusion to remove, fill, or replace objects …
657703active
chenyme/grok2api
A self-hosted multi-account API gateway for Grok Build, Grok Web, and Grok Console, written in Go with a React frontend. It exposes Grok ca…
847538active
LargeWorldModel/LWM
Large World Model (LWM) is a family of open-source 7B-parameter multimodal autoregressive transformer models trained on long videos and boo…
257425active
mediasoup
mediasoup is a powerful WebRTC SFU (Selective Forwarding Unit) implemented as a Node.js module and Rust crate, with a C++ media worker subp…
957345stable
yangchris11/samurai
SAMURAI is the official implementation of a zero-shot visual object tracker built on top of Segment Anything Model 2 (SAM 2), using a motio…
277112active
lightGallery
lightGallery is a customizable, modular, dependency-free JavaScript lightbox gallery plugin for displaying images and videos on the web and…
727049active
xushengfeng/eSearch
eSearch is a cross-platform desktop application (Electron) combining screenshot capture, offline OCR based on PaddleOCR, screen search, tra…
987036active
gaozhangmin/boxplayer
BoxPlayer is a cross-platform desktop application that unifies multiple cloud drives (Aliyun Drive, Baidu, 115, Quark, OneDrive, etc.), loc…
976878active
netless-io/flat
Flat is the open-source Web, Windows, and macOS client of Agora Flat, an online classroom platform. It provides real-time interactive white…
706415active
vllm-project/vllm-omni
vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer…
836369active
Forget-C/Jellyfish
Jellyfish is a self-hosted, end-to-end production workspace for AI-generated short dramas, covering script input, storyboard breakdown, cha…
736195active
aidlearning/AidLearning-FrameWork
AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling…
705797active
Tencent/libpag
libpag is Tencent's official real-time rendering library for PAG (Portable Animated Graphics) files, a format exported from Adobe After Eff…
985768active
DeepLabCut/DeepLabCut
DeepLabCut is an open-source Python toolbox for markerless 2D and 3D pose estimation of user-defined body parts using deep neural networks …
895745stable
Eventual-Inc/Daft
Daft is a high-performance distributed data engine with a Python dataframe API, implemented in Rust, designed for AI and multimodal workloa…
955730active
modstart-lib/aigcpanel
AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, …
885482active
Hillobar/Rope
Rope is a GUI-focused desktop application for face swapping that implements the insightface inswapper_128 model. It offers batch swapping, …
795367active
open-mmlab/mmaction2
MMAction2 is OpenMMLab's PyTorch-based toolbox and benchmark for video understanding, covering action recognition, temporal action localiza…
555142active
signalwire/freeswitch
FreeSWITCH is an open-source software-defined telecom stack that turns commodity servers into a full telephony platform for voice, video, a…
945118stable
aiortc/aiortc
aiortc is a Python library implementing WebRTC and ORTC on top of asyncio, with an API closely mirroring the JavaScript WebRTC API. It supp…
745094active
facebookresearch/AugLy
AugLy is a Python data augmentation library from Meta AI supporting audio, image, text, and video with over 100 augmentations. It focuses o…
675089stable
mediacms-io/mediacms
MediaCMS is an open source, self-hosted video and media CMS built with Django and React, exposing a REST API. It supports video, audio, ima…
995083active
facebookresearch/co-tracker
CoTracker is a transformer-based model from Meta AI and Oxford VGG that jointly tracks any point (pixel) across a video, handling occlusion…
605080active
zju3dv/EasyMocap
EasyMocap is an open-source Python toolbox for markerless human motion capture and novel view synthesis from RGB videos. It fits parametric…
544783active
Steam-Headless/docker-steam-headless
A headless Steam Docker image that runs a full Steam client and Xfce4 desktop on Linux with NVIDIA, AMD, or Intel GPU support. It enables r…
584726active
miroslavpejic85/mirotalk
MiroTalk P2P is a self-hosted, open-source WebRTC video conferencing platform that uses direct peer-to-peer connections for real-time video…
774706active
AgnesAI-Labs/AgnesAI-Models
Official gateway and model catalog for Agnes AI, providing OpenAI-compatible API access to multimodal foundation models covering text, imag…
584705active
dankamongmen/notcurses
Notcurses is a C library (with C++, Rust, and Python bindings) for building complex, vibrant textual user interfaces on modern terminal emu…
764680active
LaoFeng-mouse/flyingmouse-format
FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop…
794601active
sensity-ai/dot
dot (Deepfake Offensive Toolkit) is a Python tool that generates real-time, controllable deepfakes from a webcam feed and injects them into…
234586active
royshil/obs-backgroundremoval
An OBS Studio plugin that removes and replaces the background in portrait video using ONNX-based machine learning segmentation, acting as a…
984492active
iperov/DeepFaceLab
DeepFaceLab is the leading open-source Windows application for creating deepfakes, allowing users to swap, de-age, or replace faces and hea…
1019292maintenance
facebookresearch/jepa
Official PyTorch implementation of V-JEPA, a self-supervised method for learning visual representations from video using a joint-embedding …
294105active
hao-ai-lab/FastVideo
FastVideo is a unified Python framework for post-training and real-time inference of video diffusion models, covering data preprocessing, f…
804076active
QwenLM/Qwen2.5-Omni
Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre…
314074active
lllyasviel/Paints-UNDO
Paints-UNDO is a family of deep learning models that take an image as input and generate the step-by-step drawing sequence (sketching, inki…
404067active
umlx5h/LLPlayer
LLPlayer is a Windows media player built for language learning, featuring dual subtitles, AI-generated subtitles via Whisper ASR, real-time…
784049active
QwenLM/Qwen3-Omni
Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an…
523980active
Qsnh/meedu
MeEdu is an open-source online school (knowledge commerce) system built with PHP/Laravel, MySQL, and Redis, supporting paid video courses, …
903855active
google-research/scenic
Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr…
763821active
Perfare/AssetStudio
AssetStudio is a Windows desktop GUI tool for exploring, extracting, and exporting assets and asset bundles from Unity games and projects. …
1015604maintenance
NExT-GPT/NExT-GPT
NExT-GPT is an end-to-end any-to-any multimodal large language model that accepts and generates arbitrary combinations of text, image, vide…
373638active
Haivision/srt
SRT (Secure Reliable Transport) is an open-source transport protocol and C++ library for ultra-low-latency live video and audio streaming o…
953585stable
MooreThreads/Moore-AnimateAnyone
An open-source reproduction of AnimateAnyone that animates a character from a single reference image using pose sequences from a driving vi…
263514active
PKU-YuanGroup/Video-LLaVA
Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into…
273500active
emacs-eaf/emacs-application-framework
Emacs Application Framework (EAF) is an extensible framework that brings modern graphical capabilities to Emacs by bridging Emacs Lisp with…
773487active
Kedreamix/Linly-Talker
Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, …
483436active
roflcoopter/viseron
Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti…
983431active
XiaoMi/xiaomi-miloco
Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an…
823292active
gozfree/gear-lib
Gear-Lib is a collection of POSIX C libraries for IoT, embedded, and network service development, covering data structures, networking prot…
233223active
wendy7756/AI-Video-Transcriber
An open-source AI tool that transcribes, summarizes, and archives videos and podcasts from 30+ platforms (YouTube, TikTok, Bilibili, etc.) …
623215active
xianfei/SysMocap
SysMocap is a cross-platform, video-driven real-time motion capture system that animates 3D virtual characters from webcam footage. It rend…
783199active
facebookresearch/tribev2
TRIBE v2 is a multimodal deep learning model from Meta AI that predicts fMRI brain responses to naturalistic video, audio, and text stimuli…
553172active
miroslavpejic85/mirotalksfu
MiroTalk SFU is a self-hosted, open-source WebRTC video conferencing platform built on the Mediasoup SFU architecture, positioning itself a…
773087active
yuyuyzl/EasyVtuber
EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int…
623051active
TanStack/ai
TanStack AI is a type-safe, provider-agnostic TypeScript SDK for building AI applications with streaming chat, tool calling, agents, struct…
803028active
q191201771/lal
LAL is an audio/video live streaming broadcast server written in Go, comparable to nginx-rtmp-module but with more features. It supports RT…
233026active
SharpAI/DeepCamera
DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r…
863019active
sherlockchou86/VideoPipe
VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates …
542931active
InternLM/InternLM-XComposer
InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u…
382925active
datascale-ai/opentalking
OpenTalking is an open-source Python framework for building real-time AI digital-human (talking avatar) conversation products. It orchestra…
652897active
deforum/sd-webui-deforum
Deforum is the official extension for AUTOMATIC1111's Stable Diffusion webui that generates AI animations from text prompts using keyframed…
232859active
katelya77/KatelyaTV
KatelyaTV is a self-hosted, cross-platform video aggregation player built on Next.js 14 and TypeScript, forked from MoonTV/LunaTV. It aggre…
412857active
fqscfqj/Y2A-Auto
Y2A-Auto is a self-hosted Python application that automates re-uploading YouTube videos to AcFun and bilibili. It handles the full pipeline…
822845active
TheSmallHanCat/flow2api
Flow2API is a self-hosted Python/FastAPI service that exposes an OpenAI- and Gemini-compatible API on top of Google Flow (VideoFX/ImageFX) …
602834active
QwenLM/Qwen-MM-Plugins
A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.…
572777active
martinrotter/rssguard
RSS Guard is a cross-platform desktop feed reader supporting RSS, ATOM, JSON and many web-based feed services, with a built-in podcast/medi…
992731active
xdit-project/xDiT
xDiT is a scalable inference engine for Diffusion Transformers (DiTs) that enables parallel deployment across multiple GPUs and machines. I…
772699active
darkzOGx/youtube-automation-agent
AgentTube is a self-hosted Node.js application that uses AI agents to run a YouTube channel end to end: researching topics, writing scripts…
822666active
om-ai-lab/OmAgent
OmAgent is a Python library for building multimodal language agents, wrapping worker orchestration, task queues, and graph-based workflow o…
312665active
anliyuan/Ultralight-Digital-Human
An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It…
642627active
tl-open-source/tl-rtc-file
A self-hostable web application that uses WebRTC peer-to-peer connections to transfer files, video, screen shares, live streams, and text b…
232624active
morethanwords/tweb
Telegram Web K is the open-source TypeScript web client powering web.telegram.org/k/, based on the original Webogram and actively patched a…
772623active
kairos-agi/kairos
Kairos is the official open-source implementation of a 4B-parameter native cross-embodiment world model that unifies video understanding, f…
572598active
dreamzero0/dreamzero
DreamZero is NVIDIA's World Action Model (WAM) that jointly predicts future video and actions from a pretrained video diffusion backbone, e…
502593active
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from …
582571active
fogsightai/fogsight
Fogsight is an open-source AI agent and animation engine powered by large language models that turns abstract concepts or words into polish…
512550active

← prev page 15 / 18 next →