domain: image-processing
1843 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| dlib Dlib is a modern C++ toolkit containing machine learning algorithms, deep learning tools, computer vision, linear algebra, and general-purp… | 86 | 14431 | stable |
| PaddlePaddle/PaddleDetection PaddleDetection is an object detection toolkit built on the PaddlePaddle deep learning framework. It provides implementations of detection,… | 73 | 14389 | active |
| qubvel-org/segmentation_models.pytorch A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar… | 70 | 11706 | stable |
| facebookresearch/dinov3 Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t… | 59 | 11249 | active |
| lucidrains/denoising-diffusion-pytorch A PyTorch implementation of Denoising Diffusion Probabilistic Models (DDPM) for generative image modeling, published on PyPI as denoising_d… | 87 | 10679 | active |
| IDEA-Research/GroundingDINO Official PyTorch implementation of Grounding DINO, an open-set object detector that fuses a Transformer-based DINO detector with grounded v… | 21 | 10515 | stable |
| m87-labs/moondream Moondream is an open-weight family of small, efficient vision language models (2B to 9B MoE) that perform image captioning, visual question… | 61 | 10014 | active |
| roboflow/rf-detr RF-DETR is a real-time transformer-based model architecture from Roboflow for object detection, instance segmentation, and keypoint detecti… | 87 | 9063 | active |
| MONAI MONAI is a PyTorch-based open-source framework for deep learning in healthcare imaging, providing domain-specific transforms, 3D architectu… | 87 | 8634 | stable |
| lucidrains/imagen-pytorch A PyTorch implementation of Imagen, Google's text-to-image neural network based on cascading DDPMs conditioned on T5 text embeddings. It pr… | 23 | 8424 | active |
| open-mmlab/mmpose MMPose is an open-source pose estimation toolbox and benchmark built on PyTorch as part of the OpenMMLab ecosystem. It provides implementat… | 39 | 7855 | active |
| facebookresearch/sam-3d-objects SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code… | 55 | 7322 | active |
| BVLC/caffe Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp… | 23 | 34556 | maintenance |
| CMU-Perceptual-Computing-Lab/openpose OpenPose is a real-time multi-person keypoint detection library that jointly estimates human body, face, hand, and foot keypoints (135 tota… | 23 | 34413 | maintenance |
| chenfei-wu/TaskMatrix TaskMatrix (Visual ChatGPT) is a Python framework that connects ChatGPT with a suite of visual foundation models like Stable Diffusion, Gro… | 30 | 34003 | maintenance |
| vllm-project/vllm-omni vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer… | 83 | 6369 | active |
| KevinMusgrave/pytorch-metric-learning A PyTorch library providing modular losses, miners, samplers, and testers for deep metric learning. It supports building complete train/tes… | 48 | 6339 | active |
| ByteDance-Seed/Bagel BAGEL is an open-source unified multimodal foundation model with 7B active parameters (14B total) trained on interleaved multimodal data. I… | 55 | 6159 | active |
| open-edge-platform/anomalib Anomalib is a Python deep learning library for anomaly detection, offering state-of-the-art unsupervised algorithms for detecting and local… | 98 | 6088 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| NVIDIA/DALI NVIDIA DALI is a GPU-accelerated data loading and preprocessing library with optimized building blocks and an execution engine for deep lea… | 92 | 5734 | active |
| matterport/Mask_RCNN A Python implementation of the Mask R-CNN model for object detection and instance segmentation, built on Keras and TensorFlow with a ResNet… | 23 | 25567 | maintenance |
| Deci-AI/super-gradients SuperGradients is an open-source PyTorch-based training library for building, training, and fine-tuning state-of-the-art computer vision mo… | 54 | 5052 | active |
| KaiyangZhou/deep-person-reid Torchreid is a PyTorch library for deep-learning person re-identification, supporting both image and video reid with end-to-end training an… | 50 | 4900 | stable |
| google/wuffs Wuffs is a memory-safe programming language plus a standard library for safely parsing, decoding and encoding untrusted file formats such a… | 75 | 4820 | active |
| layumi/Person_reID_baseline_pytorch A small, friendly PyTorch baseline implementation for person and vehicle re-identification (ReID). It reproduces strong top-conference resu… | 65 | 4446 | stable |
| xlite-dev/lite.ai.toolkit A lightweight C++ toolkit providing unified APIs for 100+ pre-trained AI models across inference backends like ONNX Runtime, MNN, TensorRT,… | 74 | 4427 | active |
| bowang-lab/MedSAM MedSAM is a fine-tuned Segment Anything Model (SAM) foundation model for universal medical image segmentation, trained on over 1.5 million … | 29 | 4379 | active |
| willnorris/imageproxy A caching image proxy server written in Go that resizes, crops, and rotates remote images on the fly via URL options. It supports request s… | 73 | 3986 | stable |
| mseitzer/pytorch-fid A PyTorch port of the official TensorFlow implementation of the Fréchet Inception Distance (FID), a metric for measuring similarity between… | 23 | 3851 | stable |
| MrForExample/ComfyUI-3D-Pack An extensive ComfyUI custom node suite for processing 3D inputs like meshes and UV textures using algorithms such as 3D Gaussian Splatting … | 51 | 3850 | active |
| open-mmlab/mmpretrain MMPretrain is OpenMMLab's PyTorch-based toolbox and benchmark for image classification model pre-training, covering supervised, self-superv… | 23 | 3850 | active |
| canvg/canvg canvg is a JavaScript library that parses SVG files (from URL or text) and renders them onto HTML Canvas, including animation support. It c… | 79 | 3834 | active |
| lightly-ai/lightly LightlySSL is a Python library built on PyTorch for self-supervised learning on images, offering modular implementations of methods like Si… | 93 | 3797 | active |
| thu-ml/SageAttention SageAttention is a family of quantized attention kernels (INT8/FP8/FP4) that accelerate transformer inference 2-5x over FlashAttention with… | 40 | 3684 | active |
| JamieMason/ImageOptim-CLI A Rust CLI that automates ImageOptim, ImageAlpha, and JPEGmini on macOS to batch-optimize images as part of an automated build process. It … | 85 | 3529 | active |
| openseadragon/openseadragon OpenSeadragon is an open-source, web-based viewer for high-resolution zoomable images, implemented in pure JavaScript for desktop and mobil… | 94 | 3500 | stable |
| aleju/imgaug imgaug is a Python library for augmenting images in machine learning experiments, converting a small set of input images into a much larger… | 23 | 14741 | maintenance |
| facebookresearch/ijepa Official PyTorch implementation of I-JEPA, a self-supervised learning method that predicts latent representations of image regions from oth… | 10 | 3489 | active |
| CompVis/latent-diffusion The official research code and pretrained model zoo for Latent Diffusion Models (LDM), the paper behind Stable Diffusion, enabling high-res… | 32 | 14133 | maintenance |
| OpenGVLab/InternGPT InternGPT (iGPT) is an open-source demo platform for showcasing AI models through a pointing-language-driven visual interactive system, sup… | 29 | 3205 | active |
| facebookresearch/dinov2 PyTorch implementation and pretrained models for DINOv2, a self-supervised vision transformer method from Meta AI that learns robust visual… | 68 | 13266 | maintenance |
| cleardusk/3DDFA_V2 3DDFA_V2 is the official PyTorch implementation of the ECCV 2020 paper 'Towards Fast, Accurate and Stable 3D Dense Face Alignment'. It regr… | 23 | 3149 | stable |
| osmr/imgclsmob A research sandbox providing (re)implementations of numerous deep learning computer vision models for classification, segmentation, detecti… | 23 | 3016 | active |
| sunsmarterjie/yolov12 YOLOv12 is a PyTorch implementation of attention-centric real-time object detectors, published at NeurIPS 2025. It provides detection model… | 59 | 2947 | active |
| insidegui/AssetCatalogTinkerer A macOS application for opening Apple asset catalog (.car) files and browsing, copying, or exporting the images inside them. It also ships … | 58 | 2874 | active |
| bytedeco/javacpp-presets JavaCPP Presets provides Java bindings for commonly used native C++ libraries such as OpenCV, FFmpeg, and many others, built on the JavaCPP… | 86 | 2850 | active |
| OpenGVLab/InternImage InternImage is a large-scale CNN-based vision foundation model that uses deformable convolutions as its core operator, released with pretra… | 28 | 2841 | stable |
| lucidrains/DALLE2-pytorch A PyTorch implementation of OpenAI's DALL-E 2 text-to-image synthesis model, focusing on the diffusion prior network that predicts image em… | 23 | 11306 | maintenance |
| espressif/esp32-camera Espressif's official camera driver library for ESP32-series SoCs (ESP32, ESP32-S2, ESP32-S3), supporting a wide range of image sensors like… | 89 | 2771 | active |
| voxelmorph/voxelmorph VoxelMorph is a Python library for learning-based image registration and alignment, using unsupervised deep learning to model deformations … | 76 | 2748 | active |
| TMElyralab/MusePose MusePose is a diffusion-based, pose-guided image-to-video generation framework for creating virtual human videos, where a character in a re… | 28 | 2701 | active |
| JIA-Lab-research/LISA LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati… | 31 | 2674 | active |
| OmniSVG/OmniSVG OmniSVG is a family of end-to-end multimodal SVG generation models built on pre-trained Vision-Language Models, released with inference cod… | 51 | 2590 | active |
| Tencent-Hunyuan/HY-World-2.0 HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from … | 58 | 2571 | active |
| sthalles/SimCLR A PyTorch reference implementation of SimCLR, a self-supervised contrastive learning framework for learning visual representations from unl… | 23 | 2491 | stable |
| TorchIO-project/torchio TorchIO is a Python library for loading, augmenting, and processing 3D medical images (MRI, CT) within PyTorch deep learning pipelines. It … | 93 | 2439 | active |
| roboflow/inference Roboflow Inference is a Python library and self-hostable inference server for deploying computer vision models on any computer or edge devi… | 91 | 2427 | active |
| ailia-ai/ailia-models A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,… | 77 | 2385 | active |
| CoinCheung/pytorch-loss A PyTorch library providing a collection of loss functions (focal loss, triplet loss, AMSoftmax, label-smooth CE, dice loss, lovasz-softmax… | 32 | 2252 | active |
| apple/ml-ferret Apple's Ferret, an end-to-end multimodal large language model (MLLM) that accepts any-form referring and grounds anything in its responses,… | 27 | 8674 | maintenance |
| ellisdg/3DUnetCNN A PyTorch library for building, training, and applying 3D U-Net convolutional neural networks for medical image segmentation. It provides c… | 45 | 2224 | active |
| LiheYoung/Depth-Anything Depth Anything is a monocular depth estimation foundation model trained on 1.5M labeled and 62M+ unlabeled images, released as a Python lib… | 26 | 8195 | maintenance |
| vitoplantamura/OnnxStream A lightweight C++ inference library for ONNX models that streams weights to run large models in very little memory, accelerated by XNNPACK.… | 59 | 2086 | active |
| alibaba/EasyCV EasyCV is an all-in-one PyTorch-based computer vision toolkit from Alibaba covering self-supervised learning, vision transformers, and majo… | 32 | 1954 | active |
| eriklindernoren/PyTorch-YOLOv3 A minimal PyTorch implementation of YOLOv3 supporting training, inference, and evaluation, with compatibility for YOLOv4 and YOLOv7 weights… | 32 | 7440 | maintenance |
| Yuliang-Liu/Monkey Monkey is a large multi-modal model (LMM) research project from CVPR 2024 that improves image understanding via higher input resolution and… | 65 | 1951 | active |
| f0ng/captcha-killer-modified A modified version of the captcha-killer Burp Suite extension that intercepts captcha images from HTTP responses and recognizes them using … | 41 | 1948 | active |
| Code-with-Beto/snapai SnapAI is a Node.js CLI that generates 1024x1024 mobile app icons and 1024x500 Google Play feature graphics using OpenAI or Google Gemini i… | 76 | 1923 | active |
| visual-layer/fastdup fastdup is a free Python tool for rapidly analyzing image and video datasets to surface duplicates, outliers, broken, dark, bright, blurry,… | 67 | 1904 | active |
| lucidrains/byol-pytorch A PyTorch library implementing the Bootstrap Your Own Latent (BYOL) self-supervised learning method from DeepMind. It wraps any image-based… | 58 | 1903 | active |
| qqwweee/keras-yolo3 A Keras (TensorFlow backend) implementation of YOLOv3 for object detection, including Darknet weight conversion, image/video detection scri… | 32 | 7114 | maintenance |
| laugh12321/TensorRT-YOLO A C++/Python deployment toolkit for running YOLO-family models (YOLOv3 through YOLO26) on NVIDIA GPUs using TensorRT, with custom plugins, … | 63 | 1880 | active |
| 188080501/JQTools JQTools is an open-source developer toolbox built with Qt/QML/C++ that bundles common small utilities: text processing, hash and encryption… | 90 | 1840 | active |
| apple/ml-4m 4M is a framework from Apple and EPFL for training any-to-any multimodal foundation models using masked modeling over discrete tokens acros… | 35 | 1808 | active |
| Emu Series Emu3 is a suite of state-of-the-art multimodal models from BAAI trained solely with next-token prediction, tokenizing images, text, and vid… | 57 | 1778 | active |
| nsfw-filter/nsfw-filter A free, open-source, privacy-focused browser extension that blocks NSFW images using on-device AI classification with TensorFlow.js. It hid… | 85 | 1775 | active |
| AIdea AIdea is a fully open-source cross-platform mobile and desktop app built with Flutter that integrates mainstream large language models (GPT… | 51 | 1762 | active |
| tkarras/progressive_growing_of_gans Official TensorFlow implementation of the ICLR 2018 NVIDIA paper 'Progressive Growing of GANs', which trains generators and discriminators … | 32 | 6179 | maintenance |
| Gen-Verse/MMaDA MMaDA is an open-source family of multimodal large diffusion language models that unify textual reasoning, multimodal understanding, and te… | 49 | 1668 | active |
| thtrieu/darkflow Darkflow is a Python library that translates Darknet's YOLO neural network definitions to TensorFlow, enabling real-time object detection a… | 32 | 6139 | maintenance |
| dmlc/gluon-cv GluonCV is a deep learning toolkit providing state-of-the-art computer vision model implementations with 170+ pre-trained models. It suppor… | 23 | 5916 | maintenance |
| ml4a/ml4a ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li… | 32 | 1602 | active |
| Tencent-Hunyuan/HunyuanWorld-Voyager HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ… | 52 | 1590 | active |
| Drexubery/ViewCrafter ViewCrafter is a research codebase that uses video diffusion models to synthesize high-fidelity novel views of scenes from a single or spar… | 49 | 1587 | active |
| yakhyo/uniface UniFace is a unified Python library for face analysis that bundles detection, recognition, landmark localization, face parsing, gaze estima… | 88 | 1586 | active |
| BloodAxe/pytorch-toolbelt A Python library of PyTorch extensions providing building blocks for fast R&D prototyping, including encoder-decoder architectures, special… | 44 | 1574 | active |
| microsoft/Mage Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage… | 57 | 1516 | active |
| NVlabs/describe-anything Describe Anything Model (DAM) is a vision-language model that generates detailed descriptions of user-specified regions in images and video… | 32 | 1514 | active |
| yunjey/stargan Official PyTorch implementation of StarGAN, a unified generative adversarial network for multi-domain image-to-image translation (CVPR 2018… | 32 | 5296 | maintenance |
| FeiYull/TensorRT-Alpha A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl… | 32 | 1460 | active |
| dbolya/yolact YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++… | 51 | 5241 | maintenance |
| zylo117/Yet-Another-EfficientDet-Pytorch A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim… | 23 | 5238 | maintenance |
| amdegroot/ssd.pytorch A PyTorch implementation of the Single Shot MultiBox Detector (SSD) object detection model from the 2016 paper by Wei Liu et al. It include… | 32 | 5221 | maintenance |
| 98.js JS Paint is a web-based, pixel-perfect remake of classic MS Paint with modern extras like themes, more file formats, touch support, and acc… | 43 | 1423 | active |
| Zejun-Yang/AniPortrait AniPortrait is a Python framework from Tencent that generates photorealistic portrait animations from an audio clip and a reference image, … | 25 | 5021 | maintenance |
| huggingface/finetrainers finetrainers is a Hugging Face library for scalable, memory-optimized training (fine-tuning) of diffusion models, including LoRA training o… | 62 | 1358 | active |
| bytedance/Lance Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained … | 55 | 1329 | active |
| open-edge-platform/geti Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru… | 98 | 1317 | active |
| huawei-noah/Efficient-Computing A collection of efficient deep learning methods from Huawei Noah's Ark Lab, covering model compression, knowledge distillation, pruning, qu… | 32 | 1307 | active |