Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: video-processing

1714 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
X-PLUG/mPLUG-Owl
mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and…
362539active
microsoft/ResearchStudio
ResearchStudio is a Microsoft collection of AI agent skills that cover the entire research lifecycle, from an under-specified research dire…
592523active
dino/dino
Dino is a modern open-source XMPP (Jabber) chat client for Linux desktops, built with GTK4 and Vala. It supports end-to-end encryption via …
762480active
synctv-org/synctv
SyncTV is a self-hosted Rust server for real-time synchronized video watching with rooms, chat, and livestreaming (RTMP/HLS/HTTP-FLV/RTSP).…
942474active
SickChill/sickchill
SickChill is a self-hosted, cross-platform automatic video library manager (PVR) for TV shows written in Python with a web interface. It wa…
672442active
volcengine/ai-app-lab
AI App Lab from Volcano Engine (Volcengine) provides Arkitect, a high-code Python SDK for building LLM applications, plus Demohouse, a coll…
832433active
roboflow/inference
Roboflow Inference is a Python library and self-hostable inference server for deploying computer vision models on any computer or edge devi…
912427active
snapotter-hq/SnapOtter
SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities…
802362active
facebookresearch/perception_models
Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan…
542353active
Xilinx/PYNQ
PYNQ is an open-source Python framework from AMD/Xilinx for designing embedded systems on Zynq and other adaptive computing platforms (FPGA…
782339active
stephengpope/no-code-architects-toolkit
A self-hostable Flask-based API that consolidates common media processing tasks—video editing, captioning, audio conversion, transcription,…
502339active
jeremyckahn/chitchatter
Chitchatter is a free, open-source communication tool offering secure peer-to-peer chat, video, audio, and file sharing directly between br…
752321active
samuelgursky/davinci-resolve-mcp
A Model Context Protocol (MCP) server that lets AI assistants like Claude control DaVinci Resolve Studio through its official Scripting API…
832300active
kerwincui/FastBee
FastBee is a lightweight, full-stack open-source IoT platform built on Spring Boot with a built-in Netty MQTT broker, device management, th…
662275active
OlafenwaMoses/ImageAI
ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob…
238877maintenance
hkchengrex/MMAudio
MMAudio is a PyTorch-based model for generating synchronized audio from video and/or text inputs, using multimodal joint training across au…
432264active
nv-tlabs/lyra
Project Lyra is NVIDIA's open series of generative 3D world models, including Lyra 1.0 for feed-forward 3D/4D scene generation from a singl…
592262active
rushindrasinha/youtube-shorts-pipeline
Verticals v3 (repo youtube-shorts-pipeline) is a Python CLI that automates producing and publishing YouTube Shorts: it researches a topic, …
612251active
Alpha-VLLM/Lumina-T2X
Lumina-T2X is a unified framework for text-to-any-modality generation built on flow-based large diffusion transformers. It supports generat…
282250active
andreknieriem/open-headunit
Open Headunit is an Android app that turns an Android tablet, phone, or aftermarket head unit into an Android Auto receiver. It is a revive…
792206active
learnhouse/learnhouse
LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com…
992203active
baresip/baresip
Baresip is a portable and modular SIP User-Agent written in C with audio and video call support. It provides a rich feature set including m…
952202active
liangdabiao/Seedance2-Storyboard-Generator
A Claude Code skill and workflow toolkit that converts novels or stories into multi-episode video series by generating four-act screenplays…
552200active
TransWithAI/Faster-Whisper-TransWithAI-ChickenRice
A high-performance audio/video transcription and translation application built on Faster Whisper, optimized for Japanese-to-Chinese transla…
772196active
OpenVidu/openvidu
OpenVidu is a self-hosted, open-source video conferencing platform and WebRTC SDK suite, built as a LiveKit fork with mediasoup as its SFU …
912126active
liballeg/allegro5
Allegro 5 is a cross-platform C library for video game and multimedia programming, handling windows, input, graphics, audio, fonts, and vid…
842120stable
ozgrozer/ai-renamer
A Node.js CLI tool that uses local AI models (via Ollama or LM Studio) or OpenAI to intelligently rename files based on their contents, inc…
262111active
MiniMax-AI/cli
The official CLI for the MiniMax AI Platform, written in TypeScript, that generates text, images, video, speech, and music from the termina…
812073active
ossia/score
ossia score is a free, open-source interactive sequencer for audio-visual artists, designed to create interactive shows, installations, and…
922053active
oil-oil/oil-motion
Oil Motion is an agent-agnostic Skill that designs, generates, and integrates interactive web animations from AI-generated video. It handle…
572041active
stackia/rtp2httpd
A lightweight IPTV streaming relay server written in C that converts multicast RTP/UDP, RTSP, and HLS sources into unicast HTTP streams. It…
902013active
ppwwyyxx/wechat-dump
A Python-based tool that extracts and parses WeChat message history from a rooted Android phone, decoding the local message database and me…
542010active
akiver/cs-demo-manager
A free and open-source companion application for managing and analyzing Counter-Strike (CS2/CSGO) demo files. It extracts statistics, gener…
971988active
open-mmlab/mmagic
MMagic is OpenMMLab's toolbox for generative and multimodal AI image/video creation, built on PyTorch. It provides a large model zoo coveri…
237457maintenance
sipsorcery-org/sipsorcery
SIPSorcery is a C#/.NET library implementing SIP, WebRTC, RTP, ICE, STUN and SDP for building real-time communication applications. It is p…
771930active
TypesettingTools/Aegisub
Aegisub is a free, cross-platform advanced subtitle editor for creating and modifying subtitles. It provides audio-based timing, powerful s…
791928active
lanyeeee/bilibili-video-downloader
A cross-platform GUI desktop application built with Tauri for downloading videos, audio, subtitles, danmaku, and covers from Bilibili. It s…
751912active
diffgram/diffgram
Diffgram is a self-hosted AI datastore for managing schemas, BLOBs, and predictions, with built-in human supervision (data labeling), data …
621909active
TheSpaghettiDetective/obico-server
Obico Server is the self-hostable backend of the Obico smart 3D printing platform, providing AI-based print failure detection, webcam strea…
761899active
szczyglis-dev/py-gpt
PyGPT is an open-source, all-in-one desktop AI assistant for Linux, Windows, and Mac, written in Python. It supports chat, agents, vision, …
921892active
lanbinleo/bili2text
bili2text is a Python command-line tool that converts Bilibili videos into text transcripts from a link or BV number, handling download, au…
561873active
MixLabPro/comfyui-mixlab-nodes
A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec…
671863active
thygate/stable-diffusion-webui-depthmap-script
An extension for AUTOMATIC1111's Stable Diffusion WebUI that generates high-resolution depth maps from images using models like Marigold, M…
321853active
RanFeng/clipsketch-ai
ClipSketch AI is a web-based AI content creation workbench that imports videos from Bilibili and Xiaohongshu links, lets users frame-accura…
431840active
boundless-large-model/boundless-world-model
Boundless-World-Model (BWM) is a physically consistent, action-conditioned video world model built on Wan2.2-TI2V-5B that acts as a low-cos…
591824active
martinlaxenaire/curtainsjs
curtains.js is a lightweight vanilla WebGL JavaScript library that converts HTML DOM elements containing images, videos, and canvases into …
391824active
whotto/Video_note_generator
A Python tool that converts video URLs into polished Xiaohongshu (Little Red Book) notes and blog articles. It downloads videos, transcribe…
431808active
Emu Series
Emu3 is a suite of state-of-the-art multimodal models from BAAI trained solely with next-token prediction, tokenizing images, text, and vid…
571778active
githubXiaowangzi/NP-Manager
NP-Manager is an Android application for APK, DEX, JAR, Smali, PDF, and media file manipulation. It provides reverse-engineering features s…
761758active
ThioJoe/Auto-Synced-Translated-Dubs
A Python CLI tool that automatically translates video subtitles into multiple languages and generates AI voice dubbed audio tracks synced t…
591747active
NVIDIA-NeMo/Curator
NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for …
861736active
skalesapp/skales
Skales is a personal AI agent desktop and mobile application that runs locally on Windows, macOS, Linux, Android, and iOS, executing multi-…
821727active
jau123/MeiGen-AI-Design-MCP
An open-source MCP server that adds AI image and video generation capabilities to AI coding tools like Claude Code, Cursor, and Codex. It s…
801719active
Novage/p2p-media-loader
P2P Media Loader is an open-source JavaScript/TypeScript library that enables peer-to-peer delivery of live and on-demand HLS and MPEG-DASH…
981712active
tddworks/baguette
Baguette is a Swift CLI and WebSocket server that provides headless control of iOS simulators without opening Xcode or Simulator.app. It su…
811704active
bytedance/Sa2VA
Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understa…
701666active
XueZeyue/DanceGRPO
Official implementation of DanceGRPO, a framework applying Group Relative Policy Optimization (GRPO) to fine-tune visual generation models …
401648active
JIA-Lab-research/ControlNeXt
ControlNeXt is the official implementation of a controllable generation method for images and videos, built on Stable Diffusion XL, Stable …
241646active
Benexl/yt-x
A POSIX-compliant shell script that lets you browse YouTube and other yt-dlp-supported sites from the terminal using fzf or from an app lau…
851645active
InterDigitalInc/CompressAI
CompressAI is a PyTorch library and evaluation platform for end-to-end learned data compression research. It provides custom layers, entrop…
731627active
Lynpoint/CyberVerse
CyberVerse is an open-source, self-hosted framework for building real-time, voice-first AI agents with optional digital-human video (talkin…
631612active
ml4a/ml4a
ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li…
321602active
X-LANCE/AniTalker
AniTalker is the official PyTorch implementation of an ACM MM 2024 paper that animates a single static portrait into a vivid talking-face v…
241598active
Drexubery/ViewCrafter
ViewCrafter is a research codebase that uses video diffusion models to synthesize high-fidelity novel views of scenes from a single or spar…
491587active
finnvoor/yap
yap is a Swift CLI for on-device speech transcription of audio and video files using Apple's Speech.framework on macOS 26. It also supports…
811586active
calzoneman/sync
CyTube is a Node.js server and JavaScript/HTML web client for synchronizing online media playback across viewers in shared channels. Each c…
511581active
Tencent/DepthCrafter
DepthCrafter is a diffusion-based video depth estimation model from Tencent AI Lab that generates temporally consistent long depth sequence…
381574active
MiniMax-AI/MiniMax-MCP
The official MiniMax Model Context Protocol (MCP) server, written in Python, exposing MiniMax's text-to-speech, image generation, and video…
641569active
jpush/aurora-imui
Aurora IMUI is a general-purpose instant messaging UI component library providing MessageList and InputView components, independent of any …
235696maintenance
shrimbly/node-banana
Node Banana is an open-source, node-based visual workflow editor for building AI media generation pipelines. Users connect nodes on an infi…
811550active
sepfy/libpeer
libpeer is a portable WebRTC implementation written in C using BSD sockets, designed for IoT and embedded devices such as ESP32 and Raspber…
771542active
roncoo/roncoo-education
Roncoo Education (领课教育系统) is an open-source online education platform built with a Spring Cloud Alibaba microservices backend and Vue 3/Nux…
681539active
tin2tin/Pallaidium
Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE…
751520active
microsoft/Mage
Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage…
571516active
NVlabs/describe-anything
Describe Anything Model (DAM) is a vision-language model that generates detailed descriptions of user-specified regions in images and video…
321514active
bytedance/SALMONN
SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video…
731513active
Jamailar/Beav
Beav (formerly RedBox) is a local-first AI content operations workbench for social media creators, combining a desktop app and a Chrome/Edg…
811509active
Taiizor/Sucrose
Sucrose is a free, open-source wallpaper engine for Windows that renders interactive live wallpapers from GIFs, videos, URLs, web pages, Yo…
921503active
TIGER-AI-Lab/TheoremExplainAgent
An agentic AI system that generates long-form (5+ minute) Manim animation videos explaining mathematical and STEM theorems using LLM agents…
351499active
technomancer702/nodecast-tv
nodecast-tv is a self-hosted web-based IPTV player that streams Live TV, Movies, and Series from Xtream Codes or M3U providers directly in …
581480active
code100x/cms
An open-source content management system (LMS) that powers app.100xdevs.com, an online learning platform for coding cohorts and bootcamps. …
681461active
p2r3/beheader
A command-line tool that generates polyglot files - single files that are simultaneously valid images, videos, PDFs, ZIP archives, and HTML…
481455active
ByteDance-Seed/m3-agent
M3-Agent is a multimodal agent framework from ByteDance Seed that processes real-time visual and auditory inputs to build entity-centric lo…
481445active
valentinfrlch/ha-llmvision
LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live …
901440active
Walter0807/MotionBERT
Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske…
651439active
krusemediallc/arcads-claude-code
An agent skill pack and prompting library for generating AI marketing videos and images through the Arcads API, designed for use with Claud…
551421active
jitsi/lib-jitsi-meet
lib-jitsi-meet is a low-level JavaScript API library for building fully custom video conferencing experiences on top of Jitsi Meet infrastr…
951419active
R3gm/SoniTranslate
SoniTranslate is a Gradio-based web application that automatically dubs videos into other languages. It transcribes speech, translates it, …
551410active
Arthi-chaud/Meelo
Meelo is a self-hosted music streaming server designed for music collectors, similar to Plex or Jellyfin but focused on music. It offers ri…
991398active
zackees/transcribe-anything
A Python CLI app that transcribes local audio/video files or URLs (YouTube, Rumble, etc.) using multiple Whisper backends with automatic de…
931393active
gerbera/gerbera
Gerbera is a free, open-source UPnP/DLNA media server that streams digital media across a home network to compatible devices like TVs, game…
861390active
zelon88/HRConvert2
HRConvert2 is a self-hosted, resource-aware file conversion server written in PHP that supports 488 file formats across documents, images, …
991364active
ClimbSnail/HoloCubic_AIO
HoloCubic_AIO is an open-source all-in-one third-party firmware for the HoloCubic ESP32-based holographic desktop device, bundling apps lik…
291354active
yuanyuanxiang/SimpleRemoter
SimpleRemoter (YAMA) is a C++ remote control suite derived from the Gh0st RAT codebase, providing remote desktop, file transfer, terminal, …
601347active
wenqsun/DimensionX
DimensionX is a research framework that generates photorealistic 3D and 4D scenes from a single image using controllable video diffusion mo…
431333active
Nativ
Nativ is a free, MIT-licensed macOS desktop application for running open AI models locally on Apple Silicon Macs, built on MLX-VLM. It prov…
801330active
bytedance/Lance
Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained …
551329active
LLaVA-VL/LLaVA-NeXT
LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision …
644716maintenance
copperspice/copperspice
CopperSpice is a set of cross-platform C++ libraries (Core, Gui, Network, Multimedia, SQL, OpenGL, Vulkan, WebKit, XML, and more) derived f…
741324active
yukkcat/gemini-business2api
A self-hosted gateway service that exposes Gemini Business through an OpenAI-compatible API, with multi-account load balancing and an admin…
581321active

← prev page 16 / 18 next →