domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| XPixelGroup/DiffBIR DiffBIR is a blind image restoration framework that uses generative diffusion priors to restore degraded real-world images. It provides pre… | 34 | 4119 | maintenance |
| mlfoundations/open_flamingo OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima… | 23 | 4118 | maintenance |
| DAMO-NLP-SG/VideoLLaMA3 VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de… | 37 | 1179 | active |
| ubicomplab/rPPG-Toolbox rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and… | 51 | 1178 | active |
| NVlabs/Deep_Object_Pose NVIDIA's Deep Object Pose Estimation (DOPE), a deep learning system for detecting known objects and estimating their 6-DoF pose from RGB ca… | 48 | 1178 | active |
| balancap/SSD-Tensorflow A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net… | 32 | 4101 | maintenance |
| Soul-AILab/SoulX-LiveAct SoulX-LiveAct is the official inference code for a real-time human animation framework that generates lifelike, audio/multimodal-controlled… | 54 | 1176 | active |
| tjiiv-cprg/EPro-PnP EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc… | 41 | 1175 | stable |
| BnanZ0/ok-nte ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated… | 78 | 1174 | active |
| warmshao/FasterLivePortrait A real-time portrait animation application based on LivePortrait that animates still photos or videos using a driving video, image, audio, … | 38 | 1174 | active |
| alembic/alembic Alembic is an open computer graphics interchange framework for storing and sharing baked, animated 3D scene data, consisting of a C++ libra… | 84 | 1173 | stable |
| MIC-DKFZ/batchgenerators A Python framework for data augmentation of 2D and 3D images, developed by the German Cancer Research Center for medical image classificati… | 71 | 1173 | stable |
| chongzhou96/EdgeSAM EdgeSAM is the official PyTorch implementation of a distilled, accelerated variant of the Segment Anything Model (SAM) designed for on-devi… | 37 | 1173 | active |
| NVlabs/imaginaire NVIDIA's PyTorch library containing optimized implementations of image and video synthesis methods, including GAN-based image-to-image tran… | 32 | 4082 | maintenance |
| Kiri-Innovation/3dgs-render-blender-addon A free, open-source Blender add-on for importing, editing, animating, and rendering 3D Gaussian Splats (3DGS) inside Blender. Created by KI… | 90 | 1172 | active |
| IFL-CAMP/easy_handeye A ROS package with a GUI that automates hand-eye calibration between a robot and a camera/tracking system, supporting eye-in-hand and eye-o… | 47 | 1172 | active |
| csguoh/MambaIR MambaIR and MambaIRv2 are PyTorch-based image restoration models built on Mamba state-space models, published at ECCV 2024 and CVPR 2025. T… | 54 | 1171 | active |
| magicleap/SuperGluePretrainedNetwork SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t… | 32 | 4072 | maintenance |
| juliansteenbakker/mobile_scanner A Flutter plugin for scanning barcodes and QR codes using the device camera, backed by CameraX/ML Kit on Android, AVFoundation/Apple Vision… | 98 | 1170 | active |
| jbikker/tinybvh TinyBVH is a single-header, dependency-free C++ library for fast Bounding Volume Hierarchy (BVH) construction and ray traversal on CPU and … | 67 | 1170 | active |
| FoundationVision/GLEE GLEE is a general object foundation model for images and videos that unifies detection, segmentation, tracking, grounding, and open-world o… | 26 | 1170 | active |
| madmann91/bvh A header-only C++20 library for constructing and traversing bounding volume hierarchies (BVH), with multiple SAH-based builders, a reinsert… | 69 | 1169 | active |
| jenly1314/MLKit MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin… | 79 | 1168 | active |
| Object Detection Metrics A Python toolkit implementing the most popular metrics (AP, mAP, precision-recall curves) used to evaluate object detection algorithms, wit… | 50 | 1166 | stable |
| GaParmar/clean-fid Clean-FID is a PyTorch library for computing the Frechet Inception Distance (FID) with correct image resizing and quantization steps, fixin… | 48 | 1166 | stable |
| cure-lab/MagicDrive MagicDrive is the official PyTorch implementation of an ICLR 2024 paper for controllable street view generation using diffusion models with… | 35 | 1166 | active |
| facebookresearch/VideoPose3D A PyTorch implementation of CVPR 2019 research on 3D human pose estimation in video using temporal convolutions over 2D keypoint trajectori… | 10 | 4052 | maintenance |
| yosinski/deep-visualization-toolbox A GUI toolbox for visualizing and understanding deep neural networks, showing per-unit activations, backprop/deconv, and regularized-optimi… | 32 | 4051 | maintenance |
| CyberAgentAILab/TANGO TANGO is a research library from CyberAgent AI Lab that generates co-speech gesture videos by reenactment, using hierarchical audio-motion … | 39 | 1163 | active |
| 3D ResNets for Action Recognition A PyTorch implementation of 3D ResNet and R(2+1)D models for video action recognition, accompanying CVPR 2018 and related papers. It includ… | 23 | 4038 | maintenance |
| Linaom1214/TensorRT-For-YOLO-Series A Python and C++ toolkit for running YOLO-series object detection models (YOLOv3 through YOLOv12, YOLOX) with NVIDIA TensorRT, including ON… | 44 | 1162 | active |
| MCG-NKU/E2FGVI E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It… | 32 | 1161 | stable |
| minivision-ai/photo2cartoon A Python deep-learning project from Minivision that converts real portrait photos into cartoon-style avatars using unpaired image translati… | 32 | 4029 | maintenance |
| DepthAnything/PromptDA Prompt Depth Anything is a Python library implementing a CVPR 2025 method for high-resolution (up to 4K) accurate metric depth estimation. … | 50 | 1159 | active |
| libuvc/libuvc libuvc is a cross-platform C library for accessing USB video devices built on top of libusb. It provides fine-grained control over UVC-comp… | 32 | 1157 | active |
| fundamentalvision/Deformable-DETR Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl… | 32 | 4015 | maintenance |
| sirius-ai/LPRNet_Pytorch A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus… | 32 | 1156 | stable |
| LazarSoft/jsqrcode A JavaScript port of the ZXing QR code scanner that decodes QR codes from images or canvas in HTML5-enabled browsers. It supports webcam-ba… | 32 | 4012 | maintenance |
| JunMa11/SegLossOdyssey A curated collection of loss functions for medical image segmentation, accompanying the 'Loss Odyssey in Medical Image Segmentation' survey… | 32 | 4007 | maintenance |
| SystemErrorWang/White-box-Cartoonization Official TensorFlow implementation of the CVPR 2020 paper 'Learning to Cartoonize Using White-box Cartoon Representations', which converts … | 61 | 4001 | maintenance |
| OpenGVLab/VisionLLM VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc… | 33 | 1153 | active |
| quark0/darts DARTS is the official PyTorch implementation of the ICLR 2019 paper 'DARTS: Differentiable Architecture Search', which performs neural arch… | 32 | 3997 | maintenance |
| DennisLiu1993/Fastest_Image_Pattern_Matching A C++ library implementing an accelerated Normalized Cross Correlation (NCC)-based template matching and image alignment algorithm, based o… | 61 | 1151 | active |
| mangdangroboticsclub/QuadrupedRobot Mini Pupper is an open-source ROS-based quadruped robot dog kit built around Raspberry Pi, with software for SLAM, navigation, and OpenCV-b… | 67 | 1150 | active |
| naurril/SUSTechPOINTS SUSTechPOINTS is a web-based 3D point cloud annotation platform for labeling LiDAR data with 3D bounding boxes, aimed at autonomous driving… | 62 | 1150 | active |
| amazon-science/mm-cot Official PyTorch implementation of the paper 'Multimodal Chain-of-Thought Reasoning in Language Models', which adds vision features to a tw… | 31 | 3985 | maintenance |
| keith2018/SoftGLRender A tiny C++ software renderer/rasterizer that emulates the GPU rendering pipeline (vertex/fragment shading, rasterization, depth testing, bl… | 74 | 1149 | active |
| JDAI-CV/fast-reid FastReID is a PyTorch-based research platform implementing state-of-the-art re-identification algorithms for persons, vehicles, and faces. … | 23 | 3981 | maintenance |
| ttttccxxui/DataInfra-RedactionEverything A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and pla… | 60 | 1147 | active |
| ShiqiYu/OpenGait OpenGait is a flexible and extensible Python framework for gait recognition research, providing implementations of state-of-the-art models … | 67 | 1146 | active |
| circuitvalley/USB_C_Industrial_Camera_FPGA_USB3 Open-source USB-C industrial camera project with interchangeable C-mount lens and MIPI sensor, containing PCB designs, Lattice Crosslink NX… | 76 | 1145 | active |
| IDEA-Research/Grounding-DINO-1.5-API Python examples and API client for Grounding DINO 1.5/1.6, IDEA Research's open-world (open-set) object detection model series hosted on De… | 25 | 1143 | active |
| MIT-SPARK/Hydra Hydra is a C++ system that incrementally builds hierarchical 3D Scene Graphs from sensor data in real time. It is developed by MIT SPARK as… | 68 | 1142 | active |
| HengyiWang/spann3r Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.… | 26 | 1141 | active |
| cvg/glue-factory Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point… | 69 | 1140 | active |
| YuehaiTeam/cocogoat A browser-based toolbox for Genshin Impact that performs local achievement recognition using PaddleOCR and onnxruntime, plus achievement ma… | 76 | 1139 | active |
| clovaai/deep-text-recognition-benchmark Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio… | 32 | 3942 | maintenance |
| EyeTrackVR/EyeTrackVR EyeTrackVR is a free, open-source, DIY software platform that turns affordable cameras and IR LEDs mounted inside a VR headset into an eye … | 87 | 1138 | active |
| chengtan9907/OpenSTL OpenSTL is a comprehensive benchmark and modular framework for spatio-temporal predictive learning, covering video prediction methods acros… | 54 | 1137 | active |
| facebookresearch/watermark-anything Official PyTorch implementation and pretrained models for the paper 'Watermark Anything with Localized Messages', which embeds multiple loc… | 10 | 1137 | active |
| kzampog/cilantro cilantro is a lean, templated C++ library for processing 3D point cloud data, offering kd-trees, normal estimation, resampling, PCA, PLY I/… | 45 | 1135 | active |
| Kiteretsu77/APISR APISR is a deep-learning based super-resolution tool that restores and enhances low-quality, low-resolution anime images and videos using t… | 37 | 1135 | active |
| SarahWeiii/CoACD CoACD is a C++ library (with Python bindings and a Unity package) that performs approximate convex decomposition of 3D triangle meshes usin… | 98 | 1134 | active |
| Udayraj123/OMRChecker OMRChecker is a Python application that reads and evaluates OMR (Optical Mark Recognition) sheets scanned via a scanner or phone camera. It… | 67 | 1134 | active |
| Janspiry/Image-Super-Resolution-via-Iterative-Refinement An unofficial PyTorch implementation of SR3 (Image Super-Resolution via Iterative Refinement), a diffusion-based model for image super-reso… | 32 | 3923 | maintenance |
| zai-org/SCAIL-2 Official implementation of SCAIL-2, an open-source model for end-to-end controlled character animation that drives character videos from re… | 58 | 1132 | active |
| mgonzs13/yolo_ros A ROS 2 wrapper for Ultralytics YOLO models (YOLOv8 through YOLO26) providing object detection, tracking, instance segmentation, human pose… | 93 | 1131 | active |
| noahcao/OC_SORT OC-SORT is a pure motion-model-based multi-object tracker for video, improving on SORT by fixing Kalman filter limitations to handle occlus… | 67 | 1131 | stable |
| yohanshin/WHAM WHAM is the official PyTorch implementation of the CVPR 2024 paper 'Reconstructing World-grounded Humans with Accurate 3D Motion'. It estim… | 26 | 1130 | active |
| open-mmlab/mmtracking MMTracking is OpenMMLab's PyTorch-based toolbox for video perception tasks, unifying video object detection, multiple object tracking, sing… | 23 | 3897 | maintenance |
| sml2h3/ddddocr-fastapi A minimal FastAPI-based REST API service wrapping the DdddOcr OCR engine, exposing endpoints for image text recognition, slide captcha matc… | 23 | 1126 | active |
| brownhci/WebGazer WebGazer.js is a JavaScript eye tracking library that uses a standard webcam to predict a user's gaze location on a web page in real time. … | 65 | 3888 | maintenance |
| XieZhiFa/IdCardOCR An Android OCR library for offline recognition of Chinese second-generation ID cards, driver's licenses, and passports. It extracts all fie… | 75 | 1125 | active |
| OpenGVLab/VideoMamba VideoMamba is a state space model (Mamba-based) architecture for efficient video understanding, released with code and pretrained models fr… | 25 | 1125 | active |
| geopavlakos/hamer HaMeR (Hand Mesh Recovery) is a transformer-based model that reconstructs 3D hand meshes from single monocular images using the MANO parame… | 56 | 1124 | active |
| xverse-engine/XScene-UEPlugin An Unreal Engine 5 plugin for real-time visualization, management, editing, and scalable hybrid rendering of 3D Gaussian Splatting models. … | 40 | 1121 | active |
| princeton-vl/RAFT-Stereo RAFT-Stereo is a PyTorch implementation of a deep learning model for stereo matching that estimates disparity maps from stereo image pairs … | 75 | 1119 | stable |
| caiyuanhao1998/MST A Python toolbox for spectral compressive imaging reconstruction that implements over 15 algorithms including MST, CST, DAUHST, BiSCI, HDNe… | 53 | 1118 | active |
| yangxue0827/RotationDetection AlphaRotate is a TensorFlow-based benchmark and toolbox for rotated (oriented) object detection, implementing detectors such as R2CNN, Reti… | 23 | 1118 | active |
| storyicon/comfyui_segment_anything A ComfyUI custom node that combines GroundingDINO and Segment Anything (SAM) to segment any element in an image using semantic text prompts… | 27 | 1113 | active |
| OpenKinect/libfreenect libfreenect is a userspace driver and library for the original Microsoft Xbox Kinect sensor, providing access to RGB and depth images, moto… | 23 | 3826 | maintenance |
| Anionex/agent-vision-toolkit A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, … | 79 | 1108 | active |
| princeton-vl/DPVO DPVO is a deep learning-based visual odometry and SLAM system that estimates camera trajectories from video or image sequences using patch-… | 32 | 1108 | active |
| THU-MIG/RepViT Official PyTorch implementation of RepViT, a family of lightweight CNNs designed by integrating efficient ViT architectural designs into Mo… | 19 | 1108 | stable |
| open-mmlab/PowerPaint PowerPaint is a versatile image inpainting model (ECCV 2024) built on diffusion models that handles text-guided object insertion, object re… | 70 | 1107 | active |
| spytensor/prepare_detection_dataset A collection of Python scripts that convert object detection datasets between common annotation formats, including CSV, LabelMe JSON, COCO,… | 66 | 1107 | active |
| JEOresearch/EyeTracker A lightweight open-source Python library for 3D eye tracking that detects and fits the pupil ellipse in eye camera video or images. It is a… | 67 | 1106 | active |
| mlivesu/cinolib CinoLib is a header-only C++ library for processing polygonal and polyhedral meshes, supporting triangle, quad, and general polygon surface… | 67 | 1106 | active |
| szymanowiczs/splatter-image Official PyTorch implementation of 'Splatter Image: Ultra-Fast Single-View 3D Reconstruction' (CVPR 2024), which uses an image-to-image net… | 26 | 1106 | active |
| google-research/multinerf Google Research's official code release for three NeRF papers: Mip-NeRF 360, Ref-NeRF, and RawNeRF, written in JAX. It trains neural radian… | 10 | 3808 | maintenance |
| cameraui/camera.ui camera.ui is a self-hosted, local-first video surveillance (NVR) platform for security cameras with live viewing, 24/7 recording, and on-de… | 100 | 1104 | active |
| MIT-SPARK/VGGT-SLAM VGGT-SLAM is a dense RGB SLAM system that performs real-time feed-forward 3D scene reconstruction, optimizing on the SL(4) manifold using t… | 59 | 1102 | active |
| zclucas/RMT RMT (RuoMengTu) is a free, open-source macro and desktop automation tool built on AutoHotkey v2. It supports recording and playing keyboard… | 88 | 1101 | active |
| yzslab/gaussian-splatting-lightning A PyTorch Lightning implementation of 3D Gaussian Splatting with many derived algorithms (Mip-Splatting, LightGaussian, 2DGS, deformable Ga… | 58 | 1101 | active |
| pageauc/speed-camera A Python3 and OpenCV application that turns a Raspberry Pi, Unix, or Windows computer with a Pi camera, USB webcam, or IP/RTSP camera into … | 53 | 1101 | active |
| superxslam/SuperOdom SuperOdometry is a lightweight C++/ROS library for LiDAR-only and LiDAR-inertial odometry and mapping, developed by CMU's AirLab. It fuses … | 52 | 1101 | active |
| naturomics/CapsNet-Tensorflow A TensorFlow implementation of CapsNet (Capsule Networks) based on Geoffrey Hinton's paper 'Dynamic Routing Between Capsules'. It supports … | 32 | 3786 | maintenance |
| WangLibo1995/GeoSeg GeoSeg is an open-source PyTorch-based semantic segmentation toolbox focused on Vision Transformers for remote sensing imagery, featuring t… | 32 | 1096 | active |
| hkchengrex/Cutie Cutie is a video object segmentation framework with object-level memory reading, a follow-up to XMem offering better consistency, robustnes… | 18 | 1095 | active |
| LujiaJin/One-Pot_Multi-Frame_Denoising Official PyTorch implementation of the One-Pot Multi-frame Denoising (OPD) method published at BMVC 2022 and extended in IJCV. It provides … | 60 | 1094 | stable |