domain: image-processing
1843 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| kerlomz/captcha_trainer A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren… | 55 | 3213 | active |
| jy0205/Pyramid-Flow Pyramid Flow is the official PyTorch implementation of a training-efficient autoregressive video generation model based on pyramidal flow m… | 22 | 3208 | active |
| prs-eth/Marigold Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus… | 52 | 3198 | active |
| GuidoBartoli/sherloq Sherloq is an open-source digital image forensic toolset providing an integrated GUI environment for analyzing images for tampering and aut… | 74 | 3193 | active |
| stepfun-ai/Step-Video-T2V Step-Video-T2V is an open-source text-to-video generation model from StepFun, released with inference code and pretrained weights (includin… | 25 | 3187 | active |
| Bionus/imgbrd-grabber Grabber is a highly customizable imageboard/booru browser and mass downloader that can fetch thousands of images from multiple booru source… | 88 | 3186 | active |
| Nerogar/OneTrainer OneTrainer is a GUI and CLI application for fine-tuning diffusion image models, supporting full fine-tuning, LoRA, and embeddings across ma… | 74 | 3184 | active |
| kijai/ComfyUI-KJNodes A collection of custom nodes for ComfyUI providing utilities, model optimization, and quality-of-life improvements for node-based AI image … | 71 | 3184 | active |
| fogleman/primitive A Go command-line tool that reproduces images using geometric primitives like triangles, ellipses, and polygons. It iteratively adds shapes… | 32 | 13185 | maintenance |
| dmtrKovalenko/odiff ODiff is a very fast pixel-by-pixel image comparison library and CLI tool written in Zig with SIMD optimizations (SSE2, AVX2, AVX512, NEON)… | 97 | 3173 | active |
| Sergio0694/ComputeSharp ComputeSharp is a .NET library that lets developers write compute and pixel shaders in C# and run them in parallel on the GPU via DirectX 1… | 73 | 3161 | stable |
| ali-vilab/VGen VGen is the official repository for a holistic video generation ecosystem built on diffusion models, including the I2VGen-XL cascaded image… | 27 | 3155 | active |
| megvii-research/NAFNet NAFNet is the official PyTorch implementation of a state-of-the-art image restoration network that removes nonlinear activation functions. … | 32 | 3148 | stable |
| Djdefrag/QualityScaler QualityScaler is a Windows GUI application that uses AI deep-learning models to upscale, enhance, and de-noise images and videos. It is wri… | 95 | 3138 | active |
| nomacs/nomacs nomacs is a free, open-source image viewer for Windows, Linux, macOS, FreeBSD, and other platforms, built with Qt and optionally OpenCV. It… | 91 | 3135 | stable |
| otiai10/gosseract gosseract is a Go package that provides OCR (Optical Character Recognition) by binding to the Tesseract C++ library via cgo. It lets Go app… | 51 | 3130 | active |
| sw33tLie/macshot Macshot is a free, open-source, native macOS screenshot and screen recording tool built with Swift and AppKit. It offers region/window capt… | 74 | 3123 | active |
| google/guetzli Guetzli is a perceptual JPEG encoder from Google that produces images 20-30% smaller than libjpeg at equivalent visual quality. It is a C++… | 10 | 12913 | maintenance |
| junyanz/CycleGAN A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I… | 32 | 12870 | maintenance |
| Hugo-Dz/spritefusion-pixel-snapper A Rust-based tool (CLI, web app, and desktop edition) that snaps messy, off-grid AI-generated pixel art back onto a clean pixel grid and qu… | 58 | 3089 | active |
| micjahn/ZXing.Net ZXing.Net is a .NET port of the Java ZXing barcode library that decodes and generates barcodes such as QR Code, Data Matrix, Aztec, EAN, UP… | 72 | 3088 | active |
| allenk/GeminiWatermarkTool A C++20 tool that detects and removes the Gemini/Nano Banana Pro and Veo watermarks from AI-generated images using calibrated reverse alpha… | 79 | 3086 | active |
| zxingify/zxingify-objc ZXingObjC is a full Objective-C port of the ZXing barcode image processing library, supporting encoding and decoding of many 1D and 2D barc… | 23 | 3075 | active |
| AnyListen/tools-ocr Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model… | 23 | 3065 | active |
| yuyuyzl/EasyVtuber EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int… | 62 | 3051 | active |
| thiagoalessio/tesseract-ocr-for-php A PHP wrapper library around the Tesseract OCR command-line binary, providing a fluent API for extracting text from images. It supports mul… | 61 | 3040 | stable |
| h2non/bimg bimg is a small, fast Go library for high-level image processing built on the libvips C library via bindings. It supports operations like r… | 24 | 3030 | stable |
| Doubiiu/DynamiCrafter DynamiCrafter is an open-source research model that animates open-domain still images into short videos using pre-trained video diffusion p… | 27 | 3007 | active |
| williamyang1991/Rerender_A_Video The official PyTorch implementation of 'Rerender A Video', a SIGGRAPH Asia 2023 zero-shot text-guided video-to-video translation framework.… | 29 | 2999 | stable |
| pharmapsychotic/clip-interrogator A Python library that combines OpenAI's CLIP and Salesforce's BLIP to reverse-engineer text prompts from images, optimized for use with tex… | 23 | 2982 | stable |
| mazzzystar/Queryable Queryable is an open-source iOS app that runs Apple's MobileCLIP (formerly OpenAI's CLIP) entirely on-device to search your photo album wit… | 62 | 2977 | active |
| wasserth/TotalSegmentator TotalSegmentator is a Python command-line tool that robustly segments over 100 anatomical structures in CT and MR images using deep learnin… | 66 | 2952 | active |
| zju3dv/LoFTR LoFTR is a detector-free local image feature matching method using Transformers, released with PyTorch inference and training code plus pre… | 32 | 2950 | stable |
| gre/react-native-view-shot A React Native library that captures a view and saves it as an image, supporting both old and new architectures (Fabric + TurboModules). It… | 94 | 2947 | active |
| sylikc/jpegview JPEGView is a fast, lean, and highly configurable image viewer and editor for Windows with a minimal GUI, supporting JPEG, PNG, WEBP, TIFF,… | 23 | 2942 | active |
| KichangKim/DeepDanbooru DeepDanbooru is a Python/TensorFlow system that estimates Danbooru-style tags for anime-style girl images using multi-label classification.… | 63 | 2937 | active |
| hero8152/Infinite-Canvas An infinite canvas desktop application for orchestrating AI image, video, and LLM generation workflows. It integrates with ComfyUI, OpenAI-… | 57 | 2934 | active |
| sherlockchou86/VideoPipe VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates … | 54 | 2931 | active |
| bghira/SimpleTuner SimpleTuner is a Python fine-tuning toolkit for image, video, and audio diffusion models built on Hugging Face Diffusers. It provides a web… | 92 | 2912 | active |
| ogkalu2/comic-translate An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language… | 90 | 2911 | active |
| Saiyan-World/goku Goku is a family of flow-based (rectified flow Transformer) foundation models for joint image and video generation, released by HKU and Byt… | 23 | 2905 | active |
| idootop/MagicMirror MagicMirror is a desktop application for instant AI face swapping in photos, built with Tauri. It runs entirely offline on standard hardwar… | 37 | 2891 | active |
| spatie/image-optimizer A PHP library that optimizes PNG, JPG, WEBP, AVIF, SVG and GIF images by running them through a chain of installed optimization binaries li… | 81 | 2876 | stable |
| minimagick/minimagick MiniMagick is a lightweight Ruby wrapper around the ImageMagick command-line tool, serving as a memory-efficient alternative to RMagick. It… | 97 | 2863 | stable |
| deforum/sd-webui-deforum Deforum is the official extension for AUTOMATIC1111's Stable Diffusion webui that generates AI animations from text prompts using keyframed… | 23 | 2859 | active |
| UX-Decoder/Semantic-SAM Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired … | 33 | 2854 | active |
| joye61/pic-smaller Pic Smaller is a free, open-source batch image compressor that runs entirely in the browser, supporting JPEG, PNG, WebP, GIF, SVG, AVIF, an… | 69 | 2852 | active |
| TMElyralab/MuseV MuseV is a diffusion-based framework for generating high-fidelity virtual human videos of infinite length using a Visual Conditioned Parall… | 25 | 2846 | active |
| PythonOT/POT POT is an open-source Python library providing a large set of differentiable solvers for optimal transport problems, including exact and re… | 89 | 2837 | stable |
| TheSmallHanCat/flow2api Flow2API is a self-hosted Python/FastAPI service that exposes an OpenAI- and Gemini-compatible API on top of Google Flow (VideoFX/ImageFX) … | 60 | 2834 | active |
| metadata-extractor A Java library (with a .NET port) for reading metadata such as Exif, IPTC, XMP, and ICC profiles from image, video, and audio files. It sup… | 87 | 2825 | stable |
| openalpr/openalpr OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete… | 23 | 11452 | maintenance |
| DominikDoom/a1111-sd-webui-tagcomplete A browser-side extension for AUTOMATIC1111's Stable Diffusion web UI that provides Booru-style tag autocompletion while typing prompts. It … | 66 | 2801 | active |
| numz/ComfyUI-SeedVR2_VideoUpscaler The official ComfyUI integration of ByteDance's SeedVR2 model for high-quality video and image upscaling, provided as custom nodes. It can … | 54 | 2786 | active |
| imanoop7/Ollama-OCR A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports… | 26 | 2780 | active |
| ValentinH/react-easy-crop react-easy-crop is a React component for cropping images and videos with drag, zoom, and rotate interactions. It returns crop dimensions in… | 97 | 2772 | active |
| ideogram-oss/ideogram4 Ideogram 4 is an open-weight text-to-image foundation model trained from scratch, with inference code and weights released in Python. It fe… | 54 | 2766 | active |
| autodistill/autodistill Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab… | 29 | 2763 | active |
| NVlabs/stylegan2 The official TensorFlow implementation of StyleGAN2, NVIDIA's improved style-based generative adversarial network for high-quality uncondit… | 32 | 11184 | maintenance |
| NVIDIA/FastPhotoStyle FastPhotoStyle is NVIDIA's official PyTorch implementation of the ECCV 2018 paper 'A Closed-form Solution to Photorealistic Image Stylizati… | 23 | 11177 | maintenance |
| kha-white/manga-ocr Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to… | 90 | 2758 | stable |
| huggingface/swift-coreml-diffusers A native SwiftUI application demonstrating how to run Stable Diffusion text-to-image generation on-device using Apple's Core ML Stable Diff… | 54 | 2756 | active |
| weserv/images weserv/images is the source code of wsrv.nl, a self-hostable image cache and resize service that manipulates images on-the-fly via URL para… | 75 | 2755 | stable |
| napari/napari napari is a fast, interactive, multi-dimensional image viewer for Python, built on top of Qt and numpy. It is designed for browsing, annota… | 95 | 2739 | active |
| ModelTC/LightX2V LightX2V is a lightweight, high-performance inference framework for image and video generation, supporting tasks like text-to-video, image-… | 64 | 2733 | active |
| Audiveris/audiveris Audiveris is an open-source Optical Music Recognition (OMR) application that transcribes scanned sheet music images into symbolic music dat… | 97 | 2727 | active |
| CVCUDA/CV-CUDA CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs… | 93 | 2718 | active |
| lengstrom/fast-style-transfer A TensorFlow implementation of fast neural style transfer that applies the style of famous paintings to photos and videos in real time. It … | 32 | 10962 | maintenance |
| lobehub/sd-webui-lobe-theme Lobe Theme is a modern, highly customizable UI theme and extension for the Stable Diffusion WebUI (AUTOMATIC1111). It provides an exquisite… | 66 | 2713 | active |
| magic-research/magic-animate MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image … | 44 | 10897 | maintenance |
| xdit-project/xDiT xDiT is a scalable inference engine for Diffusion Transformers (DiTs) that enables parallel deployment across multiple GPUs and machines. I… | 77 | 2699 | active |
| IDEA-Research/T-Rex T-Rex is the official Python API client for T-Rex2, a generic open-set object detection model that combines text and visual prompts to dete… | 48 | 2699 | active |
| ComfyUI-Easy-Use ComfyUI-Easy-Use is an efficiency-focused custom nodes integration package for ComfyUI that optimizes and combines popular nodes for faster… | 78 | 2696 | active |
| SkyworkAI/SkyReels-V1 SkyReels V1 is an open-source human-centric video foundation model with Text-to-Video and Image-to-Video variants, fine-tuned from HunyuanV… | 25 | 2696 | active |
| baaivision/EVA EVA is a family of large-scale vision foundation models from BAAI, including masked image models (EVA-01/02) and scaled CLIP models (EVA-CL… | 23 | 2691 | active |
| openai/DALL-E The official PyTorch package for the discrete VAE (dVAE) component of OpenAI's DALL·E model. It does not include the transformer that gener… | 10 | 10834 | maintenance |
| bytedance/InfiniteYou InfiniteYou (InfU) is a research framework from ByteDance for identity-preserved text-to-image generation built on Diffusion Transformers l… | 37 | 2685 | active |
| civilblur/mazanoke MAZANOKE is a self-hosted, privacy-focused image optimizer that runs entirely in the browser, compressing and converting images on-device w… | 73 | 2677 | active |
| PowerHouseMan/ComfyUI-AdvancedLivePortrait A ComfyUI custom node implementing LivePortrait for fast facial expression editing and animation with real-time preview. It can edit expres… | 23 | 2675 | active |
| IceClear/StableSR StableSR is a Python research library that leverages pre-trained Stable Diffusion priors for real-world blind image super-resolution. It pr… | 21 | 2668 | stable |
| Nutlope/roomGPT RoomGPT is an open-source Next.js web application that lets users upload a photo of a room and generate redesigned variations using the Con… | 31 | 10671 | maintenance |
| hgmzhn/manga-translator-ui A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,… | 80 | 2651 | active |
| phillipi/pix2pix The original Torch (Lua) implementation of pix2pix, a conditional GAN for image-to-image translation tasks such as synthesizing photos from… | 32 | 10652 | maintenance |
| szTheory/exifcleaner ExifCleaner is a free, open-source cross-platform desktop GUI app that strips EXIF and other metadata from images, videos, and PDFs using E… | 99 | 2644 | active |
| colour-science/colour Colour is an open-source Python library providing a comprehensive collection of colour science algorithms and datasets, including colour sp… | 74 | 2642 | active |
| black-forest-labs/flux2 Official inference repository for Black Forest Labs' FLUX.2 family of open-weight image generation and editing models. It provides minimal … | 48 | 2642 | active |
| mirari/v-viewer v-viewer is an image viewer component and directive for Vue 2 and Vue 3, built on top of viewer.js. It supports rotation, scaling, zooming,… | 62 | 2638 | stable |
| thephpleague/glide Glide is a PHP library for on-demand image manipulation exposed via a simple HTTP-based API, similar to cloud services like Imgix and Cloud… | 90 | 2632 | stable |
| swz30/Restormer Restormer is an efficient Transformer architecture for high-resolution image restoration, published as a CVPR 2022 Oral paper. It provides … | 44 | 2625 | stable |
| OpenStitching/stitching A Python package providing fast and robust image stitching to create panoramas, built on OpenCV's stitching module. It offers both a Python… | 88 | 2620 | active |
| samizdatco/skia-canvas A Node.js library implementing the HTML Canvas drawing API on top of Google's Skia graphics engine, with a Rust/N-API core. It supports bit… | 81 | 2602 | active |
| crowsonkb/k-diffusion A PyTorch library implementing Karras et al. (2022) diffusion models with enhancements like improved sampling algorithms and transformer-ba… | 53 | 2600 | active |
| luca-medeiros/lang-segment-anything A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie… | 42 | 2598 | active |
| Slicer/Slicer 3D Slicer is a free, open-source desktop platform for visualization, processing, segmentation, registration, and analysis of medical and bi… | 67 | 2595 | stable |
| advimman/lama LaMa is a PyTorch-based image inpainting model that fills large missing regions in images using fast Fourier convolutions, generalizing wel… | 34 | 10217 | maintenance |
| koishijs/novelai-bot A Koishi chatbot plugin that generates images via NovelAI, with support for SD-WebUI and Stable Horde backends. It offers model/sampler/siz… | 36 | 2550 | active |
| HiDream-ai/HiDream-I1 HiDream-I1 is an open-source 17B-parameter text-to-image generative foundation model based on a Sparse Diffusion Transformer, with full and… | 34 | 2512 | active |
| luanfujun/deep-photo-styletransfer Reference implementation of the CVPR 2017 paper 'Deep Photo Style Transfer', performing photorealistic image style transfer using Torch wit… | 32 | 9989 | maintenance |
| KohakuBlueleaf/LyCORIS LyCORIS is a Python library implementing parameter-efficient fine-tuning algorithms (LoRA/LoCon, LoHa, LoKr, IA3, DyLoRA, and more) for Sta… | 73 | 2508 | active |
| LTH14/JiT A PyTorch/GPU re-implementation of JiT (Just image Transformer), a minimalist pixel-space diffusion model for high-resolution image generat… | 42 | 2507 | active |