Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: image-processing

1843 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
kerlomz/captcha_trainer
A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren…
553213active
jy0205/Pyramid-Flow
Pyramid Flow is the official PyTorch implementation of a training-efficient autoregressive video generation model based on pyramidal flow m…
223208active
prs-eth/Marigold
Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus…
523198active
GuidoBartoli/sherloq
Sherloq is an open-source digital image forensic toolset providing an integrated GUI environment for analyzing images for tampering and aut…
743193active
stepfun-ai/Step-Video-T2V
Step-Video-T2V is an open-source text-to-video generation model from StepFun, released with inference code and pretrained weights (includin…
253187active
Bionus/imgbrd-grabber
Grabber is a highly customizable imageboard/booru browser and mass downloader that can fetch thousands of images from multiple booru source…
883186active
Nerogar/OneTrainer
OneTrainer is a GUI and CLI application for fine-tuning diffusion image models, supporting full fine-tuning, LoRA, and embeddings across ma…
743184active
kijai/ComfyUI-KJNodes
A collection of custom nodes for ComfyUI providing utilities, model optimization, and quality-of-life improvements for node-based AI image …
713184active
fogleman/primitive
A Go command-line tool that reproduces images using geometric primitives like triangles, ellipses, and polygons. It iteratively adds shapes…
3213185maintenance
dmtrKovalenko/odiff
ODiff is a very fast pixel-by-pixel image comparison library and CLI tool written in Zig with SIMD optimizations (SSE2, AVX2, AVX512, NEON)…
973173active
Sergio0694/ComputeSharp
ComputeSharp is a .NET library that lets developers write compute and pixel shaders in C# and run them in parallel on the GPU via DirectX 1…
733161stable
ali-vilab/VGen
VGen is the official repository for a holistic video generation ecosystem built on diffusion models, including the I2VGen-XL cascaded image…
273155active
megvii-research/NAFNet
NAFNet is the official PyTorch implementation of a state-of-the-art image restoration network that removes nonlinear activation functions. …
323148stable
Djdefrag/QualityScaler
QualityScaler is a Windows GUI application that uses AI deep-learning models to upscale, enhance, and de-noise images and videos. It is wri…
953138active
nomacs/nomacs
nomacs is a free, open-source image viewer for Windows, Linux, macOS, FreeBSD, and other platforms, built with Qt and optionally OpenCV. It…
913135stable
otiai10/gosseract
gosseract is a Go package that provides OCR (Optical Character Recognition) by binding to the Tesseract C++ library via cgo. It lets Go app…
513130active
sw33tLie/macshot
Macshot is a free, open-source, native macOS screenshot and screen recording tool built with Swift and AppKit. It offers region/window capt…
743123active
google/guetzli
Guetzli is a perceptual JPEG encoder from Google that produces images 20-30% smaller than libjpeg at equivalent visual quality. It is a C++…
1012913maintenance
junyanz/CycleGAN
A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I…
3212870maintenance
Hugo-Dz/spritefusion-pixel-snapper
A Rust-based tool (CLI, web app, and desktop edition) that snaps messy, off-grid AI-generated pixel art back onto a clean pixel grid and qu…
583089active
micjahn/ZXing.Net
ZXing.Net is a .NET port of the Java ZXing barcode library that decodes and generates barcodes such as QR Code, Data Matrix, Aztec, EAN, UP…
723088active
allenk/GeminiWatermarkTool
A C++20 tool that detects and removes the Gemini/Nano Banana Pro and Veo watermarks from AI-generated images using calibrated reverse alpha…
793086active
zxingify/zxingify-objc
ZXingObjC is a full Objective-C port of the ZXing barcode image processing library, supporting encoding and decoding of many 1D and 2D barc…
233075active
AnyListen/tools-ocr
Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model…
233065active
yuyuyzl/EasyVtuber
EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int…
623051active
thiagoalessio/tesseract-ocr-for-php
A PHP wrapper library around the Tesseract OCR command-line binary, providing a fluent API for extracting text from images. It supports mul…
613040stable
h2non/bimg
bimg is a small, fast Go library for high-level image processing built on the libvips C library via bindings. It supports operations like r…
243030stable
Doubiiu/DynamiCrafter
DynamiCrafter is an open-source research model that animates open-domain still images into short videos using pre-trained video diffusion p…
273007active
williamyang1991/Rerender_A_Video
The official PyTorch implementation of 'Rerender A Video', a SIGGRAPH Asia 2023 zero-shot text-guided video-to-video translation framework.…
292999stable
pharmapsychotic/clip-interrogator
A Python library that combines OpenAI's CLIP and Salesforce's BLIP to reverse-engineer text prompts from images, optimized for use with tex…
232982stable
mazzzystar/Queryable
Queryable is an open-source iOS app that runs Apple's MobileCLIP (formerly OpenAI's CLIP) entirely on-device to search your photo album wit…
622977active
wasserth/TotalSegmentator
TotalSegmentator is a Python command-line tool that robustly segments over 100 anatomical structures in CT and MR images using deep learnin…
662952active
zju3dv/LoFTR
LoFTR is a detector-free local image feature matching method using Transformers, released with PyTorch inference and training code plus pre…
322950stable
gre/react-native-view-shot
A React Native library that captures a view and saves it as an image, supporting both old and new architectures (Fabric + TurboModules). It…
942947active
sylikc/jpegview
JPEGView is a fast, lean, and highly configurable image viewer and editor for Windows with a minimal GUI, supporting JPEG, PNG, WEBP, TIFF,…
232942active
KichangKim/DeepDanbooru
DeepDanbooru is a Python/TensorFlow system that estimates Danbooru-style tags for anime-style girl images using multi-label classification.…
632937active
hero8152/Infinite-Canvas
An infinite canvas desktop application for orchestrating AI image, video, and LLM generation workflows. It integrates with ComfyUI, OpenAI-…
572934active
sherlockchou86/VideoPipe
VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates …
542931active
bghira/SimpleTuner
SimpleTuner is a Python fine-tuning toolkit for image, video, and audio diffusion models built on Hugging Face Diffusers. It provides a web…
922912active
ogkalu2/comic-translate
An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language…
902911active
Saiyan-World/goku
Goku is a family of flow-based (rectified flow Transformer) foundation models for joint image and video generation, released by HKU and Byt…
232905active
idootop/MagicMirror
MagicMirror is a desktop application for instant AI face swapping in photos, built with Tauri. It runs entirely offline on standard hardwar…
372891active
spatie/image-optimizer
A PHP library that optimizes PNG, JPG, WEBP, AVIF, SVG and GIF images by running them through a chain of installed optimization binaries li…
812876stable
minimagick/minimagick
MiniMagick is a lightweight Ruby wrapper around the ImageMagick command-line tool, serving as a memory-efficient alternative to RMagick. It…
972863stable
deforum/sd-webui-deforum
Deforum is the official extension for AUTOMATIC1111's Stable Diffusion webui that generates AI animations from text prompts using keyframed…
232859active
UX-Decoder/Semantic-SAM
Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired …
332854active
joye61/pic-smaller
Pic Smaller is a free, open-source batch image compressor that runs entirely in the browser, supporting JPEG, PNG, WebP, GIF, SVG, AVIF, an…
692852active
TMElyralab/MuseV
MuseV is a diffusion-based framework for generating high-fidelity virtual human videos of infinite length using a Visual Conditioned Parall…
252846active
PythonOT/POT
POT is an open-source Python library providing a large set of differentiable solvers for optimal transport problems, including exact and re…
892837stable
TheSmallHanCat/flow2api
Flow2API is a self-hosted Python/FastAPI service that exposes an OpenAI- and Gemini-compatible API on top of Google Flow (VideoFX/ImageFX) …
602834active
metadata-extractor
A Java library (with a .NET port) for reading metadata such as Exif, IPTC, XMP, and ICC profiles from image, video, and audio files. It sup…
872825stable
openalpr/openalpr
OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete…
2311452maintenance
DominikDoom/a1111-sd-webui-tagcomplete
A browser-side extension for AUTOMATIC1111's Stable Diffusion web UI that provides Booru-style tag autocompletion while typing prompts. It …
662801active
numz/ComfyUI-SeedVR2_VideoUpscaler
The official ComfyUI integration of ByteDance's SeedVR2 model for high-quality video and image upscaling, provided as custom nodes. It can …
542786active
imanoop7/Ollama-OCR
A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports…
262780active
ValentinH/react-easy-crop
react-easy-crop is a React component for cropping images and videos with drag, zoom, and rotate interactions. It returns crop dimensions in…
972772active
ideogram-oss/ideogram4
Ideogram 4 is an open-weight text-to-image foundation model trained from scratch, with inference code and weights released in Python. It fe…
542766active
autodistill/autodistill
Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab…
292763active
NVlabs/stylegan2
The official TensorFlow implementation of StyleGAN2, NVIDIA's improved style-based generative adversarial network for high-quality uncondit…
3211184maintenance
NVIDIA/FastPhotoStyle
FastPhotoStyle is NVIDIA's official PyTorch implementation of the ECCV 2018 paper 'A Closed-form Solution to Photorealistic Image Stylizati…
2311177maintenance
kha-white/manga-ocr
Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to…
902758stable
huggingface/swift-coreml-diffusers
A native SwiftUI application demonstrating how to run Stable Diffusion text-to-image generation on-device using Apple's Core ML Stable Diff…
542756active
weserv/images
weserv/images is the source code of wsrv.nl, a self-hostable image cache and resize service that manipulates images on-the-fly via URL para…
752755stable
napari/napari
napari is a fast, interactive, multi-dimensional image viewer for Python, built on top of Qt and numpy. It is designed for browsing, annota…
952739active
ModelTC/LightX2V
LightX2V is a lightweight, high-performance inference framework for image and video generation, supporting tasks like text-to-video, image-…
642733active
Audiveris/audiveris
Audiveris is an open-source Optical Music Recognition (OMR) application that transcribes scanned sheet music images into symbolic music dat…
972727active
CVCUDA/CV-CUDA
CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs…
932718active
lengstrom/fast-style-transfer
A TensorFlow implementation of fast neural style transfer that applies the style of famous paintings to photos and videos in real time. It …
3210962maintenance
lobehub/sd-webui-lobe-theme
Lobe Theme is a modern, highly customizable UI theme and extension for the Stable Diffusion WebUI (AUTOMATIC1111). It provides an exquisite…
662713active
magic-research/magic-animate
MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image …
4410897maintenance
xdit-project/xDiT
xDiT is a scalable inference engine for Diffusion Transformers (DiTs) that enables parallel deployment across multiple GPUs and machines. I…
772699active
IDEA-Research/T-Rex
T-Rex is the official Python API client for T-Rex2, a generic open-set object detection model that combines text and visual prompts to dete…
482699active
ComfyUI-Easy-Use
ComfyUI-Easy-Use is an efficiency-focused custom nodes integration package for ComfyUI that optimizes and combines popular nodes for faster…
782696active
SkyworkAI/SkyReels-V1
SkyReels V1 is an open-source human-centric video foundation model with Text-to-Video and Image-to-Video variants, fine-tuned from HunyuanV…
252696active
baaivision/EVA
EVA is a family of large-scale vision foundation models from BAAI, including masked image models (EVA-01/02) and scaled CLIP models (EVA-CL…
232691active
openai/DALL-E
The official PyTorch package for the discrete VAE (dVAE) component of OpenAI's DALL·E model. It does not include the transformer that gener…
1010834maintenance
bytedance/InfiniteYou
InfiniteYou (InfU) is a research framework from ByteDance for identity-preserved text-to-image generation built on Diffusion Transformers l…
372685active
civilblur/mazanoke
MAZANOKE is a self-hosted, privacy-focused image optimizer that runs entirely in the browser, compressing and converting images on-device w…
732677active
PowerHouseMan/ComfyUI-AdvancedLivePortrait
A ComfyUI custom node implementing LivePortrait for fast facial expression editing and animation with real-time preview. It can edit expres…
232675active
IceClear/StableSR
StableSR is a Python research library that leverages pre-trained Stable Diffusion priors for real-world blind image super-resolution. It pr…
212668stable
Nutlope/roomGPT
RoomGPT is an open-source Next.js web application that lets users upload a photo of a room and generate redesigned variations using the Con…
3110671maintenance
hgmzhn/manga-translator-ui
A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,…
802651active
phillipi/pix2pix
The original Torch (Lua) implementation of pix2pix, a conditional GAN for image-to-image translation tasks such as synthesizing photos from…
3210652maintenance
szTheory/exifcleaner
ExifCleaner is a free, open-source cross-platform desktop GUI app that strips EXIF and other metadata from images, videos, and PDFs using E…
992644active
colour-science/colour
Colour is an open-source Python library providing a comprehensive collection of colour science algorithms and datasets, including colour sp…
742642active
black-forest-labs/flux2
Official inference repository for Black Forest Labs' FLUX.2 family of open-weight image generation and editing models. It provides minimal …
482642active
mirari/v-viewer
v-viewer is an image viewer component and directive for Vue 2 and Vue 3, built on top of viewer.js. It supports rotation, scaling, zooming,…
622638stable
thephpleague/glide
Glide is a PHP library for on-demand image manipulation exposed via a simple HTTP-based API, similar to cloud services like Imgix and Cloud…
902632stable
swz30/Restormer
Restormer is an efficient Transformer architecture for high-resolution image restoration, published as a CVPR 2022 Oral paper. It provides …
442625stable
OpenStitching/stitching
A Python package providing fast and robust image stitching to create panoramas, built on OpenCV's stitching module. It offers both a Python…
882620active
samizdatco/skia-canvas
A Node.js library implementing the HTML Canvas drawing API on top of Google's Skia graphics engine, with a Rust/N-API core. It supports bit…
812602active
crowsonkb/k-diffusion
A PyTorch library implementing Karras et al. (2022) diffusion models with enhancements like improved sampling algorithms and transformer-ba…
532600active
luca-medeiros/lang-segment-anything
A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie…
422598active
Slicer/Slicer
3D Slicer is a free, open-source desktop platform for visualization, processing, segmentation, registration, and analysis of medical and bi…
672595stable
advimman/lama
LaMa is a PyTorch-based image inpainting model that fills large missing regions in images using fast Fourier convolutions, generalizing wel…
3410217maintenance
koishijs/novelai-bot
A Koishi chatbot plugin that generates images via NovelAI, with support for SD-WebUI and Stable Horde backends. It offers model/sampler/siz…
362550active
HiDream-ai/HiDream-I1
HiDream-I1 is an open-source 17B-parameter text-to-image generative foundation model based on a Sparse Diffusion Transformer, with full and…
342512active
luanfujun/deep-photo-styletransfer
Reference implementation of the CVPR 2017 paper 'Deep Photo Style Transfer', performing photorealistic image style transfer using Torch wit…
329989maintenance
KohakuBlueleaf/LyCORIS
LyCORIS is a Python library implementing parameter-efficient fine-tuning algorithms (LoRA/LoCon, LoHa, LoKr, IA3, DyLoRA, and more) for Sta…
732508active
LTH14/JiT
A PyTorch/GPU re-implementation of JiT (Just image Transformer), a minimalist pixel-space diffusion model for high-resolution image generat…
422507active

← prev page 4 / 19 next →