Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: image-processing

1843 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
arrayfire/arrayfire
ArrayFire is a general-purpose tensor/numerical computing library for C, C++, and Python that accelerates array operations on GPUs (CUDA, O…
574902stable
tyxsspa/AnyText
AnyText is the official implementation of a diffusion-based model for multilingual visual text generation and editing in images, accepted a…
324874active
ZzzLc0405/photo-abstract-editorial
A Codex Skill / prompt template that transforms a photo into a vertical editorial composition combining the original photo with an abstract…
574868active
neilsonnn/image-blaster
A Claude skillset that converts a single input image into a full 3D environment, including meshed 3D models (.glb/.obj), Gaussian splats (.…
514822active
aloshdenny/reverse-SynthID
A research tool that reverse-engineers Google's SynthID watermark embedded in Gemini-generated images using spectral analysis and signal pr…
664815active
UX-Decoder/Segment-Everything-Everywhere-All-At-Once
SEEM is the official PyTorch implementation of the NeurIPS 2023 paper 'Segment Everything Everywhere All at Once', a model for universal im…
204794stable
Bing-su/adetailer
ADetailer is an extension for the Stable Diffusion WebUI (A1111) that automatically detects objects such as faces and hands in generated im…
744781active
open-mmlab/mmocr
MMOCR is OpenMMLab's PyTorch-based toolbox for text detection, recognition, and key information extraction. It provides a model zoo of OCR …
234752active
esimov/pigo
Pigo is a pure Go library for fast face detection, pupil/eye localization, and facial landmark detection based on the Pixel Intensity Compa…
314728stable
Kwai-Kolors/Kolors
Kolors is a large-scale latent diffusion model for photorealistic text-to-image synthesis, trained with bilingual (Chinese and English) tex…
234615active
xyxiao001/vue-cropper
A Vue.js component plugin for cropping images in the browser, supporting Vue 2 and Vue 3. It offers rotation, zooming, fixed aspect ratios,…
554559active
spipm/Depixelization_poc
Depix is a proof-of-concept tool that recovers plaintext from pixelized screenshots by matching pixelated blocks against a rendered font se…
104551active
mcmonkeyprojects/SwarmUI
SwarmUI (formerly StableSwarmUI) is a modular, self-hosted web user interface for AI image generation, supporting models like Stable Diffus…
764499active
burhanrashid52/PhotoEditor
An Android photo editing library that lets apps add paint drawing, text, filters, emoji, and stickers to images, similar to Instagram/Faceb…
734499stable
php-imagine/Imagine
Imagine is an object-oriented image manipulation library for PHP, inspired by Python's PIL. It provides a unified API over GD2, Imagick, an…
894471stable
astriaai/headshots-starter
An open-source Next.js starter kit that generates professional AI headshots from user-uploaded selfies using Astria.ai's fine-tuning and in…
404463active
Zeejay0/gathered-scenes-zine-skill
A collection of image-generation skills (prompt packs) written for Codex that transform ordinary photos into zine-style paper artworks. It …
574440active
Nutlope/restorePhotos
A Next.js web application that restores old and blurry face photos using the GFPGAN ML model via the Replicate API. It provides a hosted se…
314427active
huawei-noah/Efficient-AI-Backbones
A collection of efficient neural network backbone architectures (GhostNet, TNT, ViG, WaveMLP, TinyNet, etc.) from Huawei Noah's Ark Lab, wi…
284418active
codeforreal1/compressO
CompressO is a free, open-source desktop app for compressing videos and images to tiny sizes entirely offline, built with Tauri, Rust, and …
864415active
imazen/imageflow
Imageflow is a high-performance, memory-safe image manipulation suite written in Rust, offering a C ABI library (libimageflow), a CLI tool …
924412active
libjpeg-turbo/libjpeg-turbo
libjpeg-turbo is a SIMD-accelerated JPEG image codec library that is API/ABI-compatible with libjpeg and generally 2-6x faster. It provides…
884401stable
iperov/DeepFaceLab
DeepFaceLab is the leading open-source Windows application for creating deepfakes, allowing users to swap, de-age, or replace faces and hea…
1019292maintenance
s1dashu/ip-as-logo-skill
A compact Agent Skill that guides AI agents to generate highly simplified, rounded, subtly neo-skeuomorphic IP mascot logos. It follows the…
574374active
VectorSpaceLab/OmniGen
OmniGen is a unified diffusion-based image generation model that produces and edits images from multi-modal prompts without auxiliary modul…
474340active
OHIF/Viewers
OHIF Viewer is an open-source, zero-footprint web-based medical imaging viewer for DICOM images, built as a configurable and extensible pro…
984309active
kohler/gifsicle
Gifsicle is a command-line tool for creating, editing, and optimizing GIF images and animations, with companion programs gifview (a viewer)…
624307stable
Tencent-Hunyuan/HunyuanDiT
Hunyuan-DiT is Tencent's open-source diffusion transformer model for text-to-image generation with fine-grained Chinese language understand…
494291active
richzhang/PerceptualSimilarity
A PyTorch library implementing the LPIPS (Learned Perceptual Image Patch Similarity) metric, which measures perceptual distance between ima…
234269stable
SysCV/sam-hq
HQ-SAM (Segment Anything in High Quality) upgrades Meta's Segment Anything Model with a learnable High-Quality Output Token for accurate ze…
484255active
ali-vilab/AnyDoor
AnyDoor is the official implementation of a diffusion-based model that teleports target objects into new scenes at user-specified locations…
284238active
anthonynsimon/bild
bild is a collection of parallel image processing algorithms written in pure Go, usable both as a Go library and as a CLI tool. It supports…
964204active
evanw/thumbhash
ThumbHash is a compact encoding of an image placeholder that can be stored inline with data and rendered while the real image loads. It is …
304197stable
oxipng/oxipng
Oxipng is a multithreaded, lossless PNG/APNG compression optimizer written in Rust. It can be used as a command-line utility or as a Rust l…
944188stable
lllyasviel/style2paints
Style2Paints is an AI-driven tool that colorizes lineart sketches, optionally guided by human hints, style reference images, and lighting. …
3218179maintenance
Fannovel16/comfyui_controlnet_aux
A collection of ComfyUI custom nodes providing ControlNet auxiliary preprocessors that generate hint images (Canny edges, lineart, depth ma…
634163active
nodeca/pica
A browser-side JavaScript library for high-quality, high-speed image resizing using canvas, web workers, WebAssembly, and createImageBitmap…
764144stable
RawTherapee/RawTherapee
RawTherapee is a free, cross-platform raw photo processing program for developing images from digital cameras, written in C++ with a GTK fr…
874134active
WebODM/WebODM
WebODM is a user-friendly, commercial-grade application for drone image processing that generates georeferenced maps, point clouds, elevati…
984116active
lllyasviel/sd-forge-layerdiffuse
A Stable Diffusion WebUI (Forge) extension implementing Layer Diffusion to generate transparent images and separate foreground/background l…
254116active
VectorSpaceLab/OmniGen2
OmniGen2 is an open-source unified multimodal generation model supporting text-to-image generation, instruction-guided image editing, and i…
524112active
dominictobias/react-image-crop
A dependency-free React component for responsive image cropping with pixel or percentage coordinates. It supports touch, keyboard accessibi…
804103active
ZhengPeng7/BiRefNet
BiRefNet is a PyTorch implementation of the CAAI AIR 2024 paper 'Bilateral Reference for High-Resolution Dichotomous Image Segmentation'. I…
654098active
lllyasviel/Paints-UNDO
Paints-UNDO is a family of deep learning models that take an image as input and generate the step-by-step drawing sequence (sketching, inki…
404067active
linebender/resvg
resvg is a fast, portable SVG rendering library written entirely in safe Rust, usable as a Rust library, C library, or CLI application for …
894034stable
pinterest/PINRemoteImage
PINRemoteImage is a thread-safe, high-performance image downloading, processing, and caching library for iOS, written in Objective-C. It de…
674025stable
cshum/imagor
imagor is a fast, secure image processing server and Go library built on libvips, exposing on-the-fly image transformations via a thumbor-c…
994012active
kyleduo/TinyPNG4Mac
A native macOS application that acts as a third-party client for the TinyPNG image compression service. Users can drag and drop images or u…
754001active
PintaProject/Pinta
Pinta is a free, open-source painting and image editing program inspired by Paint.NET 3.0, built with GTK and C#/.NET. It offers drawing to…
873976active
dlemstra/Magick.NET
Magick.NET is a .NET library that wraps the ImageMagick image manipulation engine, supporting over 100 image formats. It lets C#/VB.NET/.NE…
973973stable
city96/ComfyUI-GGUF
A ComfyUI custom node pack that adds GGUF quantization support for loading diffusion models like FLUX and Stable Diffusion 3.5 in low-bitra…
513953active
Nunchaku
Nunchaku is a high-performance inference engine for 4-bit quantized diffusion models (and LLMs) based on the SVDQuant technique from an ICL…
653937active
yuanzhongqiao/printfilm
Printfilm is a self-hostable AI short-drama (short film / motion comic) creation SaaS platform built on Next.js and Spring Boot. It provide…
723913active
silvia-odwyer/photon
Photon is a high-performance image processing library written in pure Rust that compiles to WebAssembly. It provides over 90 functions for …
733884active
rh12503/triangula
Triangula is a Go application that generates triangulated and polygonal art from images using a modified genetic algorithm. It ships as a d…
563874active
SandAI-org/MAGI-1
MAGI-1 is an open-source autoregressive video generation model from Sand.ai, released with Apache-2.0 licensed code and weights. It generat…
593772active
Hunyuan-PromptEnhancer/PromptEnhancer
PromptEnhancer is a prompt-rewriting framework from Tencent Hunyuan that uses a Chain-of-Thought rewriter trained via reinforcement learnin…
613758active
Kuingsmile/PicList
PicList is an Electron-based desktop application for uploading and managing images across cloud storage and image hosting services, built o…
943748active
IDEA-Research/Grounded-SAM-2
Grounded SAM 2 is a foundation-model pipeline that combines open-set detectors (Grounding DINO, Grounding DINO 1.5/1.6, Florence-2, DINO-X)…
373708active
xinyu1205/recognize-anything
Recognize Anything is a collection of open-source image recognition foundation models, including RAM, RAM++, and Tag2Text, that perform ima…
333708active
liustack/modlens
ModLens is a vision plugin for DeepSeek Harness (dsh) and other text-only coding agents that converts pasted images into structured JSON ev…
783700active
ant-research/MagicQuill
MagicQuill is an intelligent interactive image editing system from a CVPR 2025 paper, combining a brush-based UI with AI-powered suggestion…
463688active
microsoft/Bringing-Old-Photos-Back-to-Life
The official PyTorch implementation of 'Bringing Old Photos Back to Life' (CVPR 2020 Oral), a deep learning model that restores old photos …
2315704maintenance
FluidGroup/Brightroom
Brightroom is a Swift library for building composable, customizable photo editing UIs on iOS, backed by Core Image and Metal for GPU-accele…
903667active
shitagaki-lab/see-through
A research framework from a SIGGRAPH 2026 paper that decomposes a single anime character illustration into up to 23 fully inpainted, semant…
583641active
libjxl/libjxl
libjxl is the reference implementation of the JPEG XL (ISO/IEC 18181) image format, providing both an encoder (cjxl) and decoder (djxl) as …
963637active
cmusatyalab/openface
OpenFace is a free and open source Python and Torch implementation of face recognition based on Google's FaceNet deep neural network. It ge…
6515438maintenance
albumentations-team/albumentations
Albumentations is a fast, flexible Python image augmentation library for computer vision, supporting images, masks, bounding boxes, keypoin…
1015315maintenance
SkalskiP/make-sense
makesense.ai is a free, browser-based tool for labeling photos to prepare datasets for computer vision projects. It runs entirely client-si…
233562active
ToTheBeginning/PuLID
PuLID is the official PyTorch implementation of a NeurIPS 2024 method for inserting a specific person's identity into text-to-image generat…
403550active
CookSleep/gpt_image_playground
A web-based playground for generating and editing images via the OpenAI gpt-image-2 API, built with React, TypeScript, and Tailwind CSS. It…
773549active
AliaksandrSiarohin/first-order-model
Official PyTorch/Jupyter implementation of the First Order Motion Model for image animation (NeurIPS 2019). It animates a static source ima…
3215015maintenance
Kosinkadink/ComfyUI-AnimateDiff-Evolved
A ComfyUI custom node pack providing an improved AnimateDiff integration plus advanced 'Evolved Sampling' options for animated video genera…
703531active
cszn/KAIR
A PyTorch image restoration toolbox providing training and testing code for many restoration models including DnCNN, FFDNet, SRMD, USRNet, …
233523active
MashiroSaber03/Saber-Translator
Saber-Translator is an AI-powered manga translation application that detects speech bubbles, OCRs Japanese text, translates it, inpaints th…
783516active
MooreThreads/Moore-AnimateAnyone
An open-source reproduction of AnimateAnyone that animates a character from a single reference image using pose sequences from a driving vi…
263514active
borisdayma/dalle-mini
DALL·E Mini is a Python library and model that generates images from a text prompt, available via pip and hosted on Hugging Face Model Hub.…
2314740maintenance
guandeh17/Self-Forcing
Official implementation of Self Forcing, a training method for autoregressive video diffusion models that simulates inference during traini…
373488active
Ruben2776/PicView
PicView is a fast, free, and customizable image viewer for Windows 10/11 and macOS built with C# and Avalonia. It supports a wide range of …
993487active
POSTECH-CVLab/PyTorch-StudioGAN
PyTorch-StudioGAN is a PyTorch library providing unified implementations of representative GAN architectures (BigGAN, StyleGAN2/3, etc.) fo…
233487stable
AliceVision
Meshroom is an open-source, node-based visual programming application for building and executing data processing pipelines, best known for …
693486active
jurplel/qView
qView is a minimal, fast, cross-platform desktop image viewer built with Qt5. It prioritizes space efficiency and instant image loading whi…
593485active
vietanhdev/anylabeling
AnyLabeling is a desktop image annotation tool that combines LabelImg/Labelme-style manual labeling with AI-assisted auto-labeling. It runs…
853463active
nihui/waifu2x-ncnn-vulkan
A portable command-line tool implementing the waifu2x anime-style image upscaler and denoiser using the ncnn inference framework with the V…
633456active
viliusle/miniPaint
miniPaint is a free, open-source online image editor that runs entirely in the browser using HTML5 canvas, with no server uploads required.…
723436active
NVlabs/stylegan
The official TensorFlow implementation of StyleGAN, NVIDIA's style-based generator architecture for generative adversarial networks from th…
3214416maintenance
continue-revolution/sd-webui-animatediff
An AUTOMATIC1111 Stable Diffusion WebUI extension that integrates AnimateDiff motion modules to generate animated GIFs and videos from diff…
283425active
axe312ger/sqip
SQIP is a pluggable image converter that generates tiny SVG-based low-quality image placeholders (LQIP) for improved perceived web performa…
793417active
SilenceLove/HXPhotoPicker
A Swift library for iOS providing a customizable photo/video picker supporting Live Photos, GIFs, iCloud downloads, and built-in image and …
753411active
rgthree/rgthree-comfy
A collection of custom nodes and UI improvements for ComfyUI, the node-based Stable Diffusion interface. It adds utility nodes like seed co…
633401active
IQA-PyTorch
A pure Python/PyTorch toolbox for image quality assessment (IQA) providing GPU-accelerated reimplementations of many full-reference and no-…
823380active
deepseek-ai/DeepSeek-OCR-2
DeepSeek-OCR 2 is an open-source vision-language model and inference toolkit implementing 'Visual Causal Flow' for optical character recogn…
443379active
agentheroes/agentheroes
AgentHeroes is a self-hostable open-source application for generating AI character images, animating them into videos, and scheduling the r…
323357active
xororz/local-dream
A free, open-source Android app for running Stable Diffusion locally with Snapdragon NPU acceleration, also supporting CPU/GPU inference. I…
823346active
nihui/opencv-mobile
opencv-mobile provides minimal, prebuilt OpenCV binary packages for Android, iOS, ARM Linux, Windows, Linux, macOS, HarmonyOS, WebAssembly,…
853345active
ltdrdata/ComfyUI-Impact-Pack
A custom node pack for ComfyUI that enhances Stable Diffusion image generation through detectors, detailers, upscalers, and pipeline utilit…
653283active
Tencent-Hunyuan/HunyuanImage-3.0
HunyuanImage-3.0 is Tencent's open-source native multimodal model for text-to-image and image-to-image generation, with inference code and …
573253active
cocktailpeanut/fluxgym
FluxGym is a simple web UI for training FLUX LoRA models with low VRAM support (12GB/16GB/20GB). It combines the AI-Toolkit Gradio frontend…
653246active
Physton/sd-webui-prompt-all-in-one
A Stable Diffusion WebUI extension that enhances the prompt and negative prompt input boxes with a more intuitive interface. It adds automa…
653235active
breezedeus/Pix2Text
Pix2Text is an open-source Python tool that recognizes layouts, tables, math formulas (LaTeX), and text in images and converts them into Ma…
993227active

← prev page 3 / 19 next →