Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: image-processing

4273 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
bilibili/Index-anisora
Index-AniSora is Bilibili's open-source anime video generation model, capable of creating video shots in diverse anime styles from images, …
622505active
InternLM/HuixiangDou
HuixiangDou is an LLM-based professional knowledge assistant designed for group chat scenarios, using a three-stage pipeline of preprocess,…
462502active
aurora-develop/aurora
A Go service that exposes ChatGPT Web capabilities as an OpenAI-compatible API, including chat completions, responses, file Q&A, image gene…
902498active
sthalles/SimCLR
A PyTorch reference implementation of SimCLR, a self-supervised contrastive learning framework for learning visual representations from unl…
232493stable
ppogg/YOLOv5-Lite
YOLOv5-Lite is a lightweight object detection model family evolved from YOLOv5, with models as small as ~900KB (int8) that run 10-15+ FPS o…
232490active
wdsqjq/FengYunWeather
FengYunWeather is an open-source Android weather app written in Kotlin using MVX architecture with coroutines, OkHttp, Room, and Coil. It p…
412456active
VadimBoev/FlappyBird
A Flappy Bird clone written in pure C for Android, packaged as an APK under 100 KB using OpenGL ES 2, OpenSL ES, and Android Native Activit…
762442active
nz-m/SocialEcho
SocialEcho is a full-featured social networking platform built on the MERN stack (MongoDB, Express.js, React.js, Node.js) with automated co…
312435active
openpaperwork/paperwork
Paperwork is a personal document manager for Linux and Windows that scans, OCRs, indexes, and organizes paper documents. It provides keywor…
102432active
wenlng/go-captcha
GoCaptcha is a high-performance, modular behavioral CAPTCHA library for Go that generates interactive challenges including click, slide, dr…
692418active
snapotter-hq/SnapOtter
SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities…
802412active
X-PLUG/mPLUG-DocOwl
mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO…
392411active
ValveResourceFormat/ValveResourceFormat
Source 2 Viewer (VRF) is an open-source tool for browsing VPK archives and viewing, extracting, and decompiling Source 2 game assets such a…
952409active
ailia-ai/ailia-models
A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,…
772389active
Cicada000/VV
A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d…
352377active
tencent-ailab/V-Express
V-Express is a Python research project from Tencent AI Lab that generates talking head portrait videos from a reference image, audio, and V…
252360active
facebookresearch/perception_models
Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan…
542355active
stephengpope/no-code-architects-toolkit
A self-hostable Flask-based API that consolidates common media processing tasks—video editing, captioning, audio conversion, transcription,…
502343active
Xilinx/PYNQ
PYNQ is an open-source Python framework from AMD/Xilinx for designing embedded systems on Zynq and other adaptive computing platforms (FPGA…
782340active
ArtalkJS/Artalk
Artalk is a self-hosted, open-source comment system for blogs, websites, and web applications, with a lightweight Vanilla JS client and a G…
852329active
Lingyan000/fluxdo
FluxDO is a cross-platform third-party client for the Linux.do community forum, built with Flutter and featuring a Rust-based DNS-over-HTTP…
782326active
getopenscreen/openscreen
OpenScreen is a free, open-source desktop screen recorder and video editor for Windows, macOS, and Linux that turns raw captures into polis…
822319active
Xiangyu-CAS/xiaohongshu-ops-skill
A skill for the OpenClaw agent that turns it into a Xiaohongshu (RedNote) operations assistant, using browser automation (CDP) to analyze f…
502317active
YvanYin/Metric3D
Metric3D is the official PyTorch implementation of Metric3Dv1 and Metric3Dv2, monocular geometric foundation models that predict metric dep…
332308active
mattiasgustavsson/libs
A collection of single-file, public domain (dual MIT) C/C++ libraries covering graphics app frameworks, data structures, threading, HTTP, a…
612300active
yawiii/ComfyUI-Prompt-Assistant
A ComfyUI plugin that provides an all-in-one prompt assistant, connecting to cloud LLM/VLM APIs (Zhipu, SiliconFlow, Gemini, Baidu) and loc…
702288active
helix-toolkit/helix-toolkit
Helix Toolkit is a collection of 3D components for .NET, providing XAML/MVVM-compatible scene graphs and 3D rendering for WPF, WinUI, and A…
842280active
kevinluosl/deepbot
DeepBot is a system-level AI assistant desktop application (Electron/TypeScript) that automates enterprise and personal workflows through m…
582276active
rushindrasinha/youtube-shorts-pipeline
Verticals v3 (repo youtube-shorts-pipeline) is a Python CLI that automates producing and publishing YouTube Shorts: it researches a topic, …
612270active
aigc-apps/EasyAnimate
EasyAnimate is an end-to-end Python pipeline for high-resolution, long video and image generation based on transformer diffusion (DiT) mode…
202270active
opendatalab/DocLayout-YOLO
DocLayout-YOLO is a real-time YOLO-v10-based model for detecting document layout elements (text blocks, tables, figures, etc.) in diverse d…
292263active
iText
iText is a high-performance PDF library/SDK for Java and .NET that lets developers create, manipulate, inspect, sign, and secure PDF docume…
932255stable
azavea/raster-vision
Raster Vision is an open source Python library and low-code framework for building computer vision models on satellite, aerial, and other l…
612242active
wxyhgk/retain-pdf
RetainPDF is an open-source PDF translation tool that preserves layout, formulas, and document structure, with special support for scanned/…
772238active
864381832/xJavaFxTool
xJavaFxTool is a cross-platform desktop application built with JavaFX that bundles dozens of small developer utilities, including encoding …
532238active
learnhouse/learnhouse
LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com…
992234active
aigc-apps/VideoX-Fun
VideoX-Fun is a Python-based video generation pipeline built on Diffusion Transformer models (CogVideoX-Fun, Wan-Fun) that generates videos…
672231active
microsoft/LLaVA-Med
LLaVA-Med is a large language-and-vision assistant fine-tuned for the biomedicine domain, built on the LLaVA multimodal architecture. It su…
402231active
NVIDIA/vid2vid
A PyTorch implementation of NVIDIA's video-to-video synthesis method for generating high-resolution (e.g., 2048x1024) photorealistic videos…
328692maintenance
alangrainger/immich-public-proxy
A stateless proxy that sits in front of a self-hosted Immich instance and serves only explicitly shared photos, videos, and albums to the p…
842224active
Hubs-Foundation/hubs
Hubs is an open-source, browser-based multi-user 3D virtual world and social VR platform built with A-Frame, Three.js, and WebXR/WebRTC. It…
822214active
nmwsharp/polyscope
Polyscope is a lightweight C++/Python library and interactive viewer for 3D data such as surface meshes and point clouds. It lets you regis…
742201active
PiranhaCMS/piranha.core
Piranha CMS is a lightweight, decoupled, open-source content management system for .NET 8 built on ASP.NET Core and Entity Framework Core. …
852192active
gridsome/gridsome
Gridsome is a Vue.js-powered Jamstack framework and static site generator that builds fast, CDN-ready websites from any headless CMS, APIs,…
328468maintenance
oil-oil/oil-motion
Oil Motion is an agent-agnostic Skill that designs, generates, and integrates interactive web animations from AI-generated video. It handle…
572156active
yyfz/Pi3
Pi3 (π³) is a feed-forward neural network for visual geometry reconstruction that eliminates the need for a fixed reference view, using a p…
592141active
DepthAnything/Video-Depth-Anything
Video Depth Anything is a transformer-based monocular depth estimation model for arbitrarily long videos, built on Depth Anything V2. It pr…
412133active
ozgrozer/ai-renamer
A Node.js CLI tool that uses local AI models (via Ollama or LM Studio) or OpenAI to intelligently rename files based on their contents, inc…
262114active
ZiYang-xie/WorldGen
WorldGen is a Python library that generates full 3D scenes in seconds from text prompts or images, supporting 360-degree consistent explora…
532111active
BetaStreetOmnis/xhs_ai_publisher
A desktop application for AI-powered content creation and automated publishing on Xiaohongshu (Rednote), built with PyQt5, FastAPI, and Pla…
702071active
rememberber/MooTool
MooTool is an all-in-one desktop developer toolbox offering dozens of handy utilities in a single GUI app, including JSON formatting, times…
992043active
digitalsamba/claude-code-video-toolkit
An AI-native video production toolkit designed for Claude Code, providing skills, commands, templates, and Python tools so an AI agent can …
832041active
jsdecena/laracom
Laracom is a free, open-source e-commerce application built on the Laravel PHP framework, providing a storefront and admin panel with produ…
392038active
n00mkrad/flowframes
Flowframes is a Windows GUI application for AI-based video frame interpolation, supporting RIFE (Pytorch & NCNN), DAIN (NCNN), and FLAVR (P…
702033active
Kochava-Studios/witsy
Witsy is a cross-platform desktop AI assistant built with Electron and Vue 3 that lets users chat with many LLM providers using their own A…
692024active
CaliCastle/cali.so
The open-source codebase for Cali Castle's personal website (cali.so), built with Next.js 16, React 19, TypeScript, and Tailwind CSS v4. It…
672024active
cambrian-mllm/cambrian
Cambrian-1 is a fully open family of vision-centric multimodal large language models (MLLMs) from NYU's VISIONx group, with training and ev…
472013active
ppwwyyxx/wechat-dump
A Python-based tool that extracts and parses WeChat message history from a rooted Android phone, decoding the local message database and me…
542011active
64bit/async-openai
async-openai is a Rust library providing typed, async clients for the OpenAI API, covering chat completions, responses, embeddings, assista…
932001active
Niek/chatgpt-web
A single-page web interface for OpenAI-compatible chat APIs, built with Svelte, where users bring their own API key and chats are stored pr…
751998active
cloneofsimo/lora
A Python library for applying Low-Rank Adaptation (LoRA) to quickly fine-tune text-to-image diffusion models like Stable Diffusion. It prod…
227553maintenance
adobe-research/custom-diffusion
Custom Diffusion is a research codebase for efficiently fine-tuning text-to-image diffusion models like Stable Diffusion on a few example i…
691977stable
donmccurdy/glTF-Transform
glTF Transform is an SDK for reading, editing, and writing glTF 2.0 3D models in JavaScript and TypeScript, running on both Web and Node.js…
771959active
Yuliang-Liu/Monkey
Monkey is a large multi-modal model (LMM) research project from CVPR 2024 that improves image understanding via higher input resolution and…
651951active
f0ng/captcha-killer-modified
A modified version of the captcha-killer Burp Suite extension that intercepts captcha images from HTTP responses and recognizes them using …
411949active
szczyglis-dev/py-gpt
PyGPT is an open-source, all-in-one desktop AI assistant for Linux, Windows, and Mac, written in Python. It supports chat, agents, vision, …
921903active
qqwweee/keras-yolo3
A Keras (TensorFlow backend) implementation of YOLOv3 for object detection, including Darknet weight conversion, image/video detection scri…
327114maintenance
simonw/tools
A collection of miscellaneous HTML+JavaScript single-page tools hosted at tools.simonwillison.net, almost entirely generated with LLMs as a…
691879active
NVIDIA-AI-IOT/Lidar_AI_Solution
NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil…
721871active
MixLabPro/comfyui-mixlab-nodes
A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec…
671861active
ofdrw/ofdrw
OFDRW is an open-source Java library for reading, writing, and manipulating OFD (Open Fixed-layout Document) files, a Chinese national stan…
931859active
ButzYung/SystemAnimatorOnline
XR Animator is an AI-based full-body motion capture application that uses a single webcam with MediaPipe and TensorFlow.js to drive MMD/VRM…
961854active
Kav-K/GPTDiscord
GPTDiscord is a self-hosted Discord bot providing an all-in-one GPT interface with ChatGPT-style conversations, DALL-E image generation, AI…
621853active
GongRzhe/Office-PowerPoint-MCP-Server
A Model Context Protocol (MCP) server that lets LLM clients create, edit, and manage PowerPoint (.pptx) presentations via 32 tools built on…
101852active
ever-co/ever-demand
Ever Demand is an open-source, real-time on-demand commerce platform for building single-store shops, multi-vendor marketplaces, and on-dem…
661848active
Tencent-Hunyuan/HunyuanVideo-I2V
HunyuanVideo-I2V is Tencent's open-source image-to-video generation framework built on the HunyuanVideo diffusion model, providing PyTorch …
541840active
xfangfang/Macast
Macast is a cross-platform menu bar application that turns your computer into a DLNA Media Renderer using mpv as the playback engine. It le…
236920maintenance
microlinkhq/browserless
A Node.js library that wraps Puppeteer to provide a production-ready headless Chrome/Chromium driver with built-in screenshot, PDF generati…
951833active
martinlaxenaire/curtainsjs
curtains.js is a lightweight vanilla WebGL JavaScript library that converts HTML DOM elements containing images, videos, and canvases into …
391824active
whotto/Video_note_generator
A Python tool that converts video URLs into polished Xiaohongshu (Little Red Book) notes and blog articles. It downloads videos, transcribe…
431812active
zju3dv/4K4D
4K4D is a research implementation of a 4D point cloud representation for real-time dynamic view synthesis at up to 4K resolution, built on …
271806active
jd-opensource/JoyAI-Video-Edit
JoyAI-Video-Edit is a real-time, instruction-guided video editing system that applies natural-language edits to live or uploaded video stre…
571788active
syncfusion/flutter-widgets
Syncfusion's Flutter widgets libraries providing high-quality UI widgets and file-format packages for building rich applications for iOS, A…
851779active
Totoro97/NeuS
Official PyTorch implementation of NeuS, a neural implicit surface reconstruction method that learns SDF-based surfaces via volume renderin…
321776stable
githubXiaowangzi/NP-Manager
NP-Manager is an Android application for APK, DEX, JAR, Smali, PDF, and media file manipulation. It provides reverse-engineering features s…
761766active
GAP-LAB-CUHK-SZ/gaustudio
GauStudio is a modular PyTorch framework for 3D Gaussian Splatting (3DGS) research and development, supporting novel view synthesis, 3D rec…
491762active
elder-plinius/ST3GG
ST3GG is an all-in-one steganography toolkit that hides secret data inside images, audio, documents, and network packets using 100+ encodin…
631758active
NVIDIA-NeMo/Curator
NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for …
861751active
microsoft/DirectXTK12
DirectX Tool Kit for DirectX 12 is a C++ library of helper classes for writing Direct3D 12 code, covering sprite rendering, effects, textur…
851751stable
skalesapp/skales
Skales is a personal AI agent desktop and mobile application that runs locally on Windows, macOS, Linux, Android, and iOS, executing multi-…
821745active
kijai/ComfyUI-Florence2
A ComfyUI custom node plugin that runs Microsoft's Florence-2 vision-language model for image captioning, object detection, segmentation, a…
601743active
nolanx-ai/nolanx.ai
NolanX is an open-source multi-modal agent platform for AI filmmaking that orchestrates text, image, audio, and video models into long-runn…
521723active
fish2018/YPrompt
YPrompt is a self-hosted web application that uses AI-guided conversation to elicit user requirements and automatically generate profession…
441709active
Python3Spiders/WeiboSuperSpider
A Weibo (Chinese microblog) scraping toolbox in Python covering users, topics, and comments, with extras like image downloading, sentiment …
751705active
shubham-goel/4D-Humans
4DHumans is a Python research codebase implementing HMR 2.0, a transformer-based model for 3D human mesh recovery from single images, plus …
591674active
Gen-Verse/MMaDA
MMaDA is an open-source family of multimodal large diffusion language models that unify textual reasoning, multimodal understanding, and te…
491671active
JiuhaiChen/BLIP3o
Official implementation of the BLIP3o-Series, a unified autoregressive-plus-diffusion model for text-to-image generation and editing. It co…
441667active
elixir-nx/bumblebee
Bumblebee is an Elixir library providing pre-trained neural network models built on Axon, with integration for downloading models from Hugg…
891666active
bytedance/Sa2VA
Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understa…
701666active
thunil/TecoGAN
TecoGAN is the official source code for a temporally coherent GAN for video super-resolution, published at SIGGRAPH/ACM TOG. It includes in…
326141maintenance

← prev page 38 / 43 next →