Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: image-processing

1843 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
RexanWONG/text-behind-image
An open-source web application for creating text-behind-image designs, where text appears layered behind the subject of a photo. It is avai…
622020active
XavierXiao/Dreambooth-Stable-Diffusion
An implementation of Google's Dreambooth fine-tuning method applied to Stable Diffusion, enabling personalization of a text-to-image diffus…
327738maintenance
instantX-research/InstantStyle
InstantStyle is a framework for style-preserving text-to-image generation that disentangles style and content from reference images using f…
262018active
NVlabs/SPADE
Official PyTorch implementation of SPADE (GauGAN), a CVPR 2019 method for synthesizing photorealistic images from semantic segmentation map…
327717maintenance
tdrussell/diffusion-pipe
A Python training script for fine-tuning diffusion models (image and video generation) using DeepSpeed pipeline parallelism across multiple…
672015active
webp-sh/webp_server_go
A Go-based HTTP server that serves JPEG, PNG, BMP, GIF, SVG and other images as WebP/AVIF/JXL on the fly, without changing the original URL…
952006active
nhn/tui.image-editor
TOAST UI Image Editor is a full-featured photo image editor built on HTML5 Canvas, providing crop, flip, rotate, draw, shape, text, mask, a…
237669maintenance
AntixK/PyTorch-VAE
A collection of Variational Autoencoder (VAE) model implementations in PyTorch, including Beta-VAE, VQ-VAE, IWAE, WAE, and others, with a f…
387665maintenance
kohya-ss/musubi-tuner
Musubi Tuner is a set of Python scripts for training LoRA (Low-Rank Adaptation) adapters for video and image generation model architectures…
852002active
bytetriper/RAE
Official PyTorch implementation of 'Diffusion Transformers with Representation Autoencoders' (RAE), a two-stage image generation pipeline u…
482001active
NVIDIA/Stable-Diffusion-WebUI-TensorRT
An NVIDIA extension for the Stable Diffusion Web UI (Automatic1111) that accelerates image generation using TensorRT-optimized engines on R…
181989active
zxing-cpp/zxing-cpp
ZXing-C++ is an open-source, multi-format 1D/2D barcode image processing library written in pure C++20, ported from the Java ZXing library …
971987active
antirez/iris.c
Iris is a pure C inference pipeline that generates images from text prompts using open-weights diffusion transformer models like FLUX.2 Kle…
461983active
patrikhuber/eos
A lightweight, header-only 3D Morphable Face Model (3DMM) fitting library written in modern C++11/14, with Python bindings. It provides mod…
311980active
cloneofsimo/lora
A Python library for applying Low-Rank Adaptation (LoRA) to quickly fine-tune text-to-image diffusion models like Stable Diffusion. It prod…
227550maintenance
adobe-research/custom-diffusion
Custom Diffusion is a research codebase for efficiently fine-tuning text-to-image diffusion models like Stable Diffusion on a few example i…
691978stable
JIA-Lab-research/DreamOmni2
DreamOmni2 is the official PyTorch implementation of a CVPR 2026 Highlight model for multimodal instruction-based image editing and generat…
511978active
tandpfun/wardrobe
A self-hosted web application that detects garments in photos, extracts clean product cutouts, and generates modeled editorial previews usi…
541975active
crystian/ComfyUI-Crystools
A collection of utility custom nodes and UI extensions for ComfyUI, including real-time CPU/GPU/RAM resource monitors, progress bars, and m…
391972active
theamusing/perfectPixel
A Python library that automatically detects the optimal grid size in AI-generated pixel art images and refines them into clean, perfectly a…
451971active
siliconflow/onediff
OneDiff is an out-of-the-box acceleration library for diffusion models, providing PyTorch compilation tools and optimized GPU kernels. It i…
481964active
open-mmlab/mmagic
MMagic is OpenMMLab's toolbox for generative and multimodal AI image/video creation, built on PyTorch. It provides a large model zoo coveri…
237457maintenance
thisjam/sd-webui-oldsix-prompt
A Stable Diffusion WebUI plugin that provides a categorized Chinese prompt library so users can input prompts without knowing English. It s…
281957active
Fafa-DL/Awesome-Backbones
A PyTorch-based framework that integrates many deep learning backbone models (CNNs and vision transformers like ResNet, EfficientNet, Swin …
331953active
chn-lee-yumi/MaterialSearch
MaterialSearch is a self-hosted semantic search tool that indexes local photos and videos using a CLIP multimodal model, letting users find…
731949active
LTH14/mar
Official PyTorch implementation of MAR (Masked Autoregressive) image generation with DiffLoss, from the NeurIPS 2024 paper 'Autoregressive …
541949stable
openai/guided-diffusion
OpenAI's codebase for guided diffusion models from the paper 'Diffusion Models Beat GANs on Image Synthesis', including classifier conditio…
327419maintenance
PixArt-alpha/PixArt-sigma
PixArt-Σ is a PyTorch implementation of a diffusion transformer model for high-resolution (up to 4K) text-to-image generation, trained with…
251939active
starik222/BooruDatasetTagManager
A desktop tag editor for managing booru-style tagged image and video datasets used to train Stable Diffusion models such as LoRAs, embeddin…
761938active
wysaid/android-gpuimage-plus
A C++ and Java library for Android that applies GPU-accelerated image, camera, and video filters using OpenGL shaders. It supports rule-str…
881927active
riddleling/iOS-OCR-Server
An iOS app that turns an iPhone into a local OCR server using Apple's Vision Framework, exposing an HTTP API and web interface for image te…
741927active
Yuanshi9815/OminiControl
OminiControl is a universal control framework for Diffusion Transformer models like FLUX, supporting subject-driven and spatial control (ed…
621927active
AbdBarho/stable-diffusion-webui-docker
A Docker Compose setup that runs Stable Diffusion locally with popular web UIs like AUTOMATIC1111, ComfyUI, and InvokeAI. It packages model…
237309maintenance
pymatting/pymatting
PyMatting is a Python library for alpha matting that estimates an alpha matte from an input image and a hand-drawn trimap to extract foregr…
671914active
scaleflex/filerobot-image-editor
Filerobot Image Editor is an easy-to-integrate JavaScript image editing library for web applications, supporting resize, crop, flip, finetu…
691908active
TanShilongMario/PromptFill
Prompt Fill is a web-based structured prompt generation tool for AI painting platforms like GPT, Midjourney, and Nano Banana. It lets users…
601907active
sicxu/Deep3DFaceRecon_pytorch
A PyTorch implementation of Deep3DFaceReconstruction, a weakly-supervised CNN method for reconstructing 3D face geometry from a single imag…
321907stable
Faceplugin-ltd/Open-Source-Face-Recognition-SDK
An open-source face recognition SDK by Faceplugin providing face detection, landmark extraction, feature embedding generation, and face tem…
641903active
d8ahazard/sd_dreambooth_extension
A Stable Diffusion WebUI extension that adds DreamBooth fine-tuning capabilities, ported from Shivam Shrirao's optimized Diffusers implemen…
421887active
vt-vl-lab/3d-photo-inpainting
A Python research codebase from a CVPR 2020 paper that converts a single RGB-D image into a 3D photo using layered depth inpainting. It hal…
327093maintenance
zcpua/midjourney-api
An unofficial Node.js/TypeScript client library for interacting with the MidJourney image generation service via Discord. It wraps MidJourn…
651872active
we0091234/Chinese_license_plate_detection_recognition
A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports …
711868active
MixLabPro/comfyui-mixlab-nodes
A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec…
671863active
jingsongliujing/OnnxOCR
A lightweight multilingual OCR library rebuilt from PaddleOCR models to run on ONNXRuntime, removing the PaddlePaddle dependency for fast i…
741860active
D-Ogi/WatermarkRemover-AI
An AI-powered desktop application that detects and removes watermarks from images and videos using Microsoft's Florence-2 for detection and…
561859active
gaomingqi/Track-Anything
Track-Anything is an interactive tool for video object tracking and segmentation built on Segment Anything, XMem, and E2FGVI. Users specify…
566994maintenance
rosuH/EasyWatermark
EasyWatermark is an open-source Android app for adding text or image watermarks to photos, built to protect sensitive images from leaking o…
721853active
thygate/stable-diffusion-webui-depthmap-script
An extension for AUTOMATIC1111's Stable Diffusion WebUI that generates high-resolution depth maps from images using models like Marigold, M…
321853active
LuChengTHU/dpm-solver
Official PyTorch implementation of DPM-Solver and DPM-Solver++, fast high-order ODE solvers for diffusion probabilistic model sampling that…
321852stable
bootchk/resynthesizer
A suite of third-party plugins for GIMP implementing the Resynthesizer algorithm for texture synthesis and transfer among images. It enable…
701850active
ConsistentlyInconsistentYT/Pixeltovoxelprojector
A Python tool that projects the motion of pixels onto a voxel representation, converting 2D pixel movement into 3D voxel space. It is a pop…
421849active
ivmartel/dwv
DWV (DICOM Web Viewer) is an open source, zero-footprint JavaScript/HTML5 library for viewing and manipulating DICOM medical images in any …
941844active
NVlabs/stylegan3
Official PyTorch implementation of StyleGAN3 (Alias-Free GANs), a state-of-the-art generative adversarial network for high-fidelity image s…
326943maintenance
paintdotnet/release
The official download repository for Paint.NET, providing offline installer EXEs, portable ZIPs, and deployment MSIs for the Windows image …
781842active
YangLing0818/RPG-DiffusionMaster
Official implementation of RPG (Recaption, Plan, Generate), a training-free framework that uses multimodal LLMs as prompt recaptioners and …
281842active
Tencent-Hunyuan/HunyuanVideo-I2V
HunyuanVideo-I2V is Tencent's open-source image-to-video generation framework built on the HunyuanVideo diffusion model, providing PyTorch …
541840active
RanFeng/clipsketch-ai
ClipSketch AI is a web-based AI content creation workbench that imports videos from Bilibili and Xiaohongshu links, lets users frame-accura…
431840active
NVIDIA/pix2pixHD
PyTorch implementation of pix2pixHD, a conditional GAN method for synthesizing and manipulating high-resolution (2048x1024) photorealistic …
326923maintenance
AcademySoftwareFoundation/openexr
OpenEXR is the specification and reference C/C++ implementation of the EXR file format, the professional high-dynamic-range image storage f…
991836stable
leslievan/semi-utils
A Python CLI tool that batch-adds EXIF watermarks (camera model, lens, focal length, aperture, shutter, ISO, shooting time, brand logo) to …
491835active
AlibabaResearch/AdvancedLiterateMachinery
A collection of original OCR and document understanding models, algorithms, and benchmarks from Alibaba's Tongyi Lab, including models like…
551834active
Tavris1/ComfyUI-Easy-Install
ComfyUI-Easy-Install is a portable one-click installer for ComfyUI that bundles an EZi Desktop application, requiring no manual Python or G…
851831active
timothybrooks/instruct-pix2pix
PyTorch implementation of InstructPix2Pix, a diffusion-based model that edits images according to natural language instructions (e.g., 'tur…
316885maintenance
Zheng-Chong/CatVTON
CatVTON is a lightweight diffusion model for virtual try-on that swaps clothing onto a person image using a concatenation-based architectur…
401824active
tjko/jpegoptim
jpegoptim is a command-line utility for optimizing and compressing JPEG files, offering lossless optimization via Huffman table optimizatio…
561809active
all-in-aigc/aicover
A full-stack Next.js web application that generates AI-designed cover images using DALL-E 3. It ships with user auth (Clerk), payments (Str…
271807active
hako-mikan/sd-webui-regional-prompter
A custom script extension for AUTOMATIC1111's stable-diffusion-webui that lets users assign different prompts to different regions of a gen…
551804active
OpenImagingLab/FlashVSR
FlashVSR is a one-step diffusion-based streaming video super-resolution framework that runs at ~17 FPS for 768x1408 video on a single A100 …
611799active
jazzband/sorl-thumbnail
sorl-thumbnail is a Django app that generates and manages image thumbnails with pluggable engines (Pillow, ImageMagick, Wand, etc.) and key…
881794stable
TotallyNotChase/glitch-this
A Python library and command-line tool that applies customizable glitch effects to images and converts images into glitched GIFs. It offers…
231793stable
Kosinkadink/ComfyUI-VideoHelperSuite
A ComfyUI custom node suite providing video workflow nodes such as Load Video, Load Image Sequence, and Video Combine. It converts videos t…
721792active
Webreaper/Damselfly
Damselfly is a server-based photograph management application designed to index and search very large image collections using metadata such…
881783active
xyTom/snippai
Snippai is an AI-powered snipping tool that captures screenshots and uses AI to extract structured content such as LaTeX formulas, text, ta…
891782active
huchenlei/ComfyUI-layerdiffuse
A ComfyUI custom node plugin implementing LayerDiffuse, enabling generation of transparent images with RGBA output and foreground/backgroun…
291779active
puffinsoft/jscanify
jscanify is an open-source pure JavaScript document scanning library powered by OpenCV.js. It detects and highlights documents in images an…
731769active
pondorasti/emojis
A web application that generates custom Slack emojis from text prompts using AI image generation (SDXL emoji model) via Replicate. It is a …
571768active
Coyote-A/ultimate-upscale-for-automatic1111
An extension for the AUTOMATIC1111 Stable Diffusion web UI that upscales images to 2K/4K+ by processing them in tiled passes with diffusion…
311766stable
levihsu/OOTDiffusion
Official implementation of OOTDiffusion, a latent diffusion model for controllable virtual try-on that generates images of a person wearing…
266586maintenance
xinntao/ESRGAN
ESRGAN (Enhanced SRGAN) is a PyTorch-based image super-resolution model that won the PIRM 2018 Challenge on Perceptual Super-Resolution. Th…
326568maintenance
nguyenq/tess4j
Tess4J is a Java JNA wrapper for the Tesseract OCR API, enabling optical character recognition in Java applications. It supports TIFF, JPEG…
911757stable
giventofly/pixelit
Pixel It is a JavaScript library that converts images into pixel art on an HTML canvas, with configurable pixel scale, color palettes, and …
591751active
Stability-AI/StableCascade
Official codebase for Stable Cascade, a text-to-image generation model built on the Würstchen architecture with a highly compressed latent …
266540maintenance
CompVis/taming-transformers
The official implementation of 'Taming Transformers for High-Resolution Image Synthesis' (CVPR 2021), combining a convolutional VQGAN codeb…
326521maintenance
TencentARC/BrushNet
BrushNet is the official PyTorch implementation of an ECCV 2024 plug-and-play image inpainting model that embeds pixel-level masked image f…
251745active
kijai/ComfyUI-Florence2
A ComfyUI custom node plugin that runs Microsoft's Florence-2 vision-language model for image captioning, object detection, segmentation, a…
601742active
openai/consistency_models
Official PyTorch implementation of Consistency Models, a generative image model family from OpenAI supporting consistency distillation, con…
106486maintenance
Xiaojiu-z/EasyControl
EasyControl is the official implementation of an ICCV 2025 paper adding efficient and flexible conditional control to Diffusion Transformer…
351737active
SHI-Labs/OneFormer
OneFormer is a universal image segmentation framework (CVPR 2023) that unifies semantic, instance, and panoptic segmentation in a single tr…
321736stable
NimaNzrii/comfyui-photoshop
A Photoshop plugin that integrates ComfyUI's AI image generation directly into the Photoshop workspace. It lets users run Stable Diffusion …
711734active
mahmoodlab/CLAM
CLAM is an open-source Python toolkit for data-efficient, weakly supervised classification of whole-slide images (WSIs) in computational pa…
391728active
liuruoze/EasyPR
EasyPR is an open-source C++ library built on OpenCV for recognizing Chinese license plates in unconstrained situations, outputting plate c…
236429maintenance
facebookresearch/ConvNeXt
Official PyTorch implementation of ConvNeXt, a pure convolutional neural network architecture from the CVPR 2022 paper 'A ConvNet for the 2…
106416maintenance
jau123/MeiGen-AI-Design-MCP
An open-source MCP server that adds AI image and video generation capabilities to AI coding tools like Claude Code, Cursor, and Codex. It s…
801719active
verlab/accelerated_features
XFeat is a lightweight, fast learned keypoint detector and descriptor for local feature extraction and image matching, supporting both spar…
161716active
Donaldcwl/browser-image-compression
A JavaScript library that compresses jpeg, png, webp, and bmp images directly in the web browser by reducing resolution or storage size. It…
231714stable
liip/LiipImagineBundle
LiipImagineBundle is a Symfony bundle providing an image manipulation abstraction toolkit built on the Imagine library. It lets developers …
941712active
kha-white/mokuro
mokuro is a Python tool that performs text detection and OCR on Japanese manga pages and generates overlay files (.mokuro or HTML) enabling…
861712active
imageio/imageio
Imageio is a mature Python library for reading and writing image and video data, including animated images, volumetric data, and scientific…
921711stable
lukechilds/merge-images
A small JavaScript library that composites multiple images into one, abstracting away canvas boilerplate into a single promise-based functi…
691710stable
fish2018/YPrompt
YPrompt is a self-hosted web application that uses AI-guided conversation to elicit user requirements and automatically generate profession…
441709active

← prev page 6 / 19 next →