Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: image-processing

1843 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
peterbraden/node-opencv
Native Node.js bindings for the OpenCV computer vision library, exposing Matrices, image reading/writing, and cascades like face detection …
234384maintenance
withoutbg/withoutbg-python
A Python SDK (pip install withoutbg) for removing image backgrounds, offering a free local open-weights ONNX model and an optional paid clo…
801246active
jark006/JarkViewer
JarkViewer is a minimalist, fast image viewer for 64-bit Windows supporting a huge range of formats including AVIF, HEIC, JPEG XL, RAW came…
831245active
fpgaminer/joycaption
JoyCaption is an open, free, and uncensored image captioning Visual Language Model (VLM) with released weights and training scripts. It gen…
531244active
LTH14/fractalgen
A PyTorch implementation of Fractal Generative Models (FractalGen), enabling pixel-by-pixel high-resolution image generation. It includes p…
241244active
wfjsw/danbooru-diffusion-prompt-builder
A web application ('Danbooru Tag Supermarket') for browsing, searching, and composing Danbooru/NovelAI tag prompts for Stable Diffusion ima…
321243active
showlab/Tune-A-Video
Tune-A-Video is the official PyTorch implementation of an ICCV 2023 paper that fine-tunes pre-trained text-to-image diffusion models (like …
314364maintenance
raspberrypi/picamera2
Picamera2 is a Python library providing an interface to Raspberry Pi cameras via the libcamera stack, replacing the legacy Picamera library…
941241active
XPixelGroup/HYPIR
Official PyTorch implementation of HYPIR, a SIGGRAPH 2025 method that harnesses diffusion-yielded score priors for image restoration. It pr…
391241active
jcjohnson/fast-neural-style
A Torch (Lua) implementation of feedforward neural style transfer from the ECCV 2016 paper 'Perceptual Losses for Real-Time Style Transfer …
324359maintenance
ZHKKKe/MODNet
MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima…
324355maintenance
facebookresearch/deit
Official PyTorch repository for DeiT and related vision transformer architectures (CaiT, ResMLP, PatchConvnet, DeiT III), providing trainin…
104355maintenance
bytedance/USO
USO is ByteDance's open-source unified style- and subject-driven image generation model based on diffusion (FLUX), combining any subject wi…
361235active
Acly/comfyui-inpaint-nodes
A set of custom nodes for ComfyUI that improve image inpainting and outpainting workflows. It integrates the Fooocus inpaint model for SDXL…
641232active
TheSmallHanCat/sora2api
A self-hosted OpenAI-compatible API gateway that wraps Sora's text-to-video and image generation capabilities behind standard /v1/chat/comp…
101232active
chengzeyi/Comfy-WaveSpeed
A ComfyUI custom node plugin that acts as an all-in-one inference optimization solution for diffusion models, built around First Block Cach…
661231stable
lucidrains/deep-daze
Deep Daze is a simple command line tool for text-to-image generation that combines OpenAI's CLIP with a Siren implicit neural representatio…
234315maintenance
florestefano1975/comfyui-portrait-master
A ComfyUI custom node suite that helps AI image creators generate detailed, professional prompts for human portraits. It provides modular n…
561228active
pythongosssss/ComfyUI-WD14-Tagger
A ComfyUI custom node extension that interrogates images to extract booru-style tags using WD 1.4 tagger models (ONNX-based). It integrates…
431226active
mcmonkeyprojects/sd-dynamic-thresholding
A Stable Diffusion extension that enables using higher CFG scale values without color artifacts by clamping latents between sampling steps.…
361224active
flyimg/flyimg
Flyimg is a Dockerized, self-hosted application that resizes, crops, and compresses images on the fly via URL parameters, serving optimized…
1001223active
benrugg/AI-Render
A Blender add-on that renders AI-generated images with Stable Diffusion based on a text prompt and the user's 3D scene. It supports cloud r…
571223active
maxritter/diy-thermocam
DIY-Thermocam is an open-source, self-assembly thermal imaging camera based on the FLIR Lepton sensor and a Teensy 4.1 microcontroller, wit…
481222active
tapmodo/Jcrop
Jcrop is a JavaScript image cropping engine that lets developers add interactive crop-selection functionality to images in web applications…
324273maintenance
kijai/ComfyUI-segment-anything-2
A set of ComfyUI custom nodes that bring Meta's Segment Anything 2 (SAM2) models into ComfyUI workflows for promptable image and video segm…
431214active
DachunKai/EvTexture
Official PyTorch implementation of EvTexture and EvTexture++, event-driven video super-resolution models that use event-camera signals to e…
541207active
gali8/Tesseract-OCR-iOS
An iOS framework wrapping the Tesseract OCR engine (with Leptonica and image libraries) for use in Objective-C or Swift apps on iOS 9.0+. I…
234221maintenance
willisma/SiT
Official PyTorch implementation of Scalable Interpolant Transformers (SiT), a family of generative models built on Diffusion Transformers t…
531206active
frotms/PaddleOCR2Pytorch
A PyTorch port of PaddleOCR that lets you run PaddleOCR-trained models (detection, recognition, and document structure parsing) without the…
731205active
sammycage/lunasvg
LunaSVG is a lightweight, portable C++ library for rendering and manipulating SVG files, built on PlutoVG. It can rasterize SVG documents t…
771203active
GoogleCloudPlatform/vertex-ai-creative-studio
GenMedia Creative Studio is a web application showcasing Google Cloud's generative media APIs including Gemini Image (Nano Banana), Veo vid…
901199active
cleanlab/cleanvision
CleanVision is a Python library that automatically detects issues in image datasets, such as blurry, dark, over-exposed, or near-duplicate …
591199active
guoyingtao/Mantis
Mantis is a Swift image cropping library for iOS and Mac Catalyst offering UIKit and SwiftUI APIs with an Apple Photos-style crop experienc…
941198active
Calamari-OCR/calamari
Calamari is a Python-based OCR engine for line-based automatic text recognition, built on OCRopy and Kraken with a TensorFlow deep-learning…
741197active
LasseRafn/ui-avatars
UI Avatars is a PHP service that generates simple initial-based avatar images from names via a URL-based API, powering ui-avatars.com. It c…
761196stable
haidog-yaqub/MeanFlow
An unofficial PyTorch implementation of MeanFlow and iMF, one-step generative modeling methods based on flow matching. It provides config-d…
591196active
gumlet/php-image-resize
A PHP library for resizing, scaling, and cropping images using the GD extension. It supports saving to multiple formats, loading from files…
861194active
kornelski/dssim
DSSIM is a Rust CLI tool and library that measures perceptual (dis)similarity between PNG/JPEG images using a multi-scale variant of the SS…
681194active
lessthanoptimal/BoofCV
BoofCV is an open-source, real-time computer vision library written entirely in Java, covering image processing, camera calibration, featur…
861192active
shallowdream204/DreamClear
DreamClear is a diffusion-transformer based real-world image restoration model for high-fidelity super-resolution, published at NeurIPS 202…
271191active
sksamuel/scrimage
Scrimage is an immutable, functional JVM library for image manipulation usable from Java, Kotlin, and Scala. It supports resizing, format c…
971186active
pqpo/SmartCropper
An Android library for smart image cropping that automatically detects document borders using OpenCV (with an optional TensorFlow Lite HED …
664132maintenance
advanced-cropper/vue-advanced-cropper
A flexible Vue.js image cropping component library that lets developers build custom croppers matching any website design. It supports canv…
231185active
SHI-Labs/Neighborhood-Attention-Transformer
Official PyTorch implementation of the Neighborhood Attention Transformer (NAT/DiNAT), a family of efficient vision transformers with local…
321184stable
orpatashnik/StyleCLIP
Official implementation of StyleCLIP, a method for text-driven manipulation of StyleGAN-generated imagery using CLIP. It provides three app…
324121maintenance
XPixelGroup/DiffBIR
DiffBIR is a blind image restoration framework that uses generative diffusion priors to restore degraded real-world images. It provides pre…
344119maintenance
Tencent-Hunyuan/MixGRPO
MixGRPO is a research framework from Tencent Hunyuan implementing a mixed ODE-SDE GRPO algorithm for efficient reinforcement learning fine-…
581177active
warmshao/FasterLivePortrait
A real-time portrait animation application based on LivePortrait that animates still photos or videos using a driving video, image, audio, …
381174active
chongzhou96/EdgeSAM
EdgeSAM is the official PyTorch implementation of a distilled, accelerated variant of the Segment Anything Model (SAM) designed for on-devi…
371173active
NVlabs/imaginaire
NVIDIA's PyTorch library containing optimized implementations of image and video synthesis methods, including GAN-based image-to-image tran…
324082maintenance
bytedance/1d-tokenizer
A research repository from ByteDance containing code and pretrained model weights for 1D visual tokenizers (TiTok, TA-TiTok, FlowTok) and i…
291172active
csguoh/MambaIR
MambaIR and MambaIRv2 are PyTorch-based image restoration models built on Mamba state-space models, published at ECCV 2024 and CVPR 2025. T…
541171active
FoundationVision/GLEE
GLEE is a general object foundation model for images and videos that unifies detection, segmentation, tracking, grounding, and open-world o…
261170active
wladradchenko/wunjo.wladradchenko.ru
Wunjo CE is an open-source, locally-run AI media suite for face swap, lip sync, voice cloning, object/text/background removal, restyling, a…
701169active
CH563/shot-easy-website
ShotEasy is a free online photo and screenshot toolkit built with Astro that runs entirely in the browser using WebAssembly. It offers scre…
701162active
hako-mikan/sd-webui-lora-block-weight
A custom script/extension for AUTOMATIC1111's stable-diffusion-webui that allows per-block weight control when applying LoRA models. It sup…
431162active
MCG-NKU/E2FGVI
E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It…
321161stable
minivision-ai/photo2cartoon
A Python deep-learning project from Minivision that converts real portrait photos into cartoon-style avatars using unpaired image translati…
324029maintenance
kijai/ComfyUI-IC-Light
A ComfyUI custom node that provides a native implementation of IC-Light models for image relighting. It lets users run IC-Light workflows i…
351159active
alibaba/lumenx
LumenX is an AI-native platform for turning novel text into publishable motion comic and short drama videos. It provides a full pipeline fr…
591158active
Exiv2/exiv2
Exiv2 is a cross-platform C++ library and command-line utility for reading, writing, deleting, and modifying Exif, IPTC, XMP, and ICC image…
891157stable
sirius-ai/LPRNet_Pytorch
A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus…
321156stable
picturepan2/instagram.css
A pure CSS library that recreates all 41 Instagram photo filters using only CSS classes, with no JavaScript required. Filters are applied b…
324012maintenance
SystemErrorWang/White-box-Cartoonization
Official TensorFlow implementation of the CVPR 2020 paper 'Learning to Cartoonize Using White-box Cartoon Representations', which converts …
614001maintenance
DennisLiu1993/Fastest_Image_Pattern_Matching
A C++ library implementing an accelerated Normalized Cross Correlation (NCC)-based template matching and image alignment algorithm, based o…
611151active
Woolverine94/biniou
biniou is a self-hosted web UI for 30+ generative AI models covering image, video, audio, and text generation, built with Gradio and Huggin…
631150active
Piasy/BigImageViewer
An Android library for displaying very large images with pan and zoom while using very little memory, built on Subsampling Scale Image View…
323981maintenance
PurpleDoubleD/locally-uncensored
Locally Uncensored is a free, open-source desktop AI studio (built with Tauri/TypeScript) that bundles uncensored local chat, a coding agen…
811146active
sbs20/scanservjs
scanservjs is a self-hosted web UI frontend for SANE-compatible scanners, letting you share one or more scanners over a network from a Linu…
951145active
okooo5km/HiPixel
HiPixel is a native macOS app for AI-powered image super-resolution, built with SwiftUI and using Upscayl's AI models. It offers batch upsc…
761145active
lquesada/ComfyUI-Inpaint-CropAndStitch
A set of ComfyUI custom nodes that crop an image around a masked area before sampling and stitch the inpainted result back afterward. This …
681145active
rohitgandikota/sliders
Official implementation of Concept Sliders, LoRA adaptors that enable precise, plug-and-play control of attributes in diffusion models like…
521139active
clovaai/deep-text-recognition-benchmark
Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio…
323942maintenance
JonasKruckenberg/imagetools
Imagetools is a set of import directives for Vite that transform and optimize images at compile time, powered by sharp. It supports resizin…
671138active
facebookresearch/watermark-anything
Official PyTorch implementation and pretrained models for the paper 'Watermark Anything with Localized Messages', which embeds multiple loc…
101137active
deepskystacker/DSS
DeepSkyStacker is a freeware application for registering and stacking deep-sky astrophotography images to improve signal-to-noise ratio. Or…
931135active
CellProfiler/CellProfiler
CellProfiler is a free, open-source desktop application for quantitative analysis of biological images, letting biologists build modular im…
681135active
Kiteretsu77/APISR
APISR is a deep-learning based super-resolution tool that restores and enhances low-quality, low-resolution anime images and videos using t…
371135active
rhymes-ai/Allegro
Allegro is an open-source text-to-video generation model that produces high-quality 720p videos up to 6 seconds at 15 FPS from text prompts…
241135active
Udayraj123/OMRChecker
OMRChecker is a Python application that reads and evaluates OMR (Optical Mark Recognition) sheets scanned via a scanner or phone camera. It…
671134active
Janspiry/Image-Super-Resolution-via-Iterative-Refinement
An unofficial PyTorch implementation of SR3 (Image Super-Resolution via Iterative Refinement), a diffusion-based model for image super-reso…
323923maintenance
AlexanderVanhee/Gradia
Gradia is a GNOME desktop application for enhancing screenshots before sharing, adding backgrounds, padding, annotations, and export option…
751128active
atiilla/GeoIntel
GeoIntel is a Python tool that uses Google's Gemini API to estimate where a photo was taken through AI-powered geolocation analysis. It off…
581128active
sml2h3/ddddocr-fastapi
A minimal FastAPI-based REST API service wrapping the DdddOcr OCR engine, exposing endpoints for image text recognition, slide captcha matc…
231126active
lkwq007/stablediffusion-infinity
A web application for outpainting with Stable Diffusion on an effectively infinite canvas, built with PyScript and Gradio. It extends image…
323887maintenance
gemini-cli-extensions/nanobanana
A Gemini CLI extension that adds image generation and editing commands backed by Google's Nano Banana image models. It runs as an MCP serve…
701123active
Xilinx/Vitis_Libraries
A collection of open-source, performance-optimized accelerated libraries for the AMD Vitis Unified Software Platform, offering out-of-the-b…
871122active
1j01/textual-paint
Textual Paint is a terminal-based (TUI) image editor that recreates the classic MS Paint interface, built with the Textual framework in Pyt…
601122active
premieroctet/photoshot
Photoshot is an open-source web application that generates custom AI avatars from user-uploaded selfies using fine-tuned text-to-image mode…
223870maintenance
Faster3ck/Converseen
Converseen is a free, cross-platform batch image processor built on ImageMagick that converts, resizes, rotates, and flips images in bulk a…
971119active
dacnay816y62-hub/cinema-dna-21x9x3
A prompt-engineering skill that translates characters, settings, genres, or a one-line plot into cinematic 21:9 triptych image-generation p…
761119active
caiyuanhao1998/MST
A Python toolbox for spectral compressive imaging reconstruction that implements over 15 algorithms including MST, CST, DAUHST, BiSCI, HDNe…
531118active
openai/improved-diffusion
The official codebase for OpenAI's Improved Denoising Diffusion Probabilistic Models paper, providing a Python package for training and sam…
323844maintenance
storyicon/comfyui_segment_anything
A ComfyUI custom node that combines GroundingDINO and Segment Anything (SAM) to segment any element in an image using semantic text prompts…
271113active
jina-ai/discoart
DiscoArt is a Python library that wraps Disco Diffusion (CLIP-guided diffusion) so generative artists and developers can create AI artworks…
233826maintenance
baofff/U-ViT
U-ViT is the official PyTorch implementation of a ViT-based backbone architecture for diffusion models from the CVPR 2023 paper 'All are Wo…
321110stable
open-mmlab/PowerPaint
PowerPaint is a versatile image inpainting model (ECCV 2024) built on diffusion models that handles text-guided object insertion, object re…
701107active
ATH-MaaS/Pixelle-MCP
Pixelle MCP is an open-source omnimodal AIGC framework that converts ComfyUI workflows (local or RunningHub cloud) into MCP tools with zero…
441104active
TencentARC/T2I-Adapter
Official implementation of T2I-Adapter, lightweight adapter models that add controllable conditioning (sketch, canny, lineart, depth, pose)…
313801maintenance
zai-org/CogView4
CogView4, CogView3-Plus and CogView3 are open-source text-to-image generation models from Zhipu AI, with CogView4 being a 6B-parameter DiT-…
281101active

← prev page 9 / 19 next →