function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| senguptaumd/Background-Matting Official research code for 'Background Matting: The World is Your Green Screen' (CVPR 2020), a deep network that extracts per-pixel alpha m… | 32 | 4769 | maintenance |
| IrisRainbowNeko/genshin_auto_fish A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen… | 23 | 4758 | maintenance |
| liwenxi/SWIFT-AI SWIFT-AI is a deep learning system for extremely fast gigapixel-level visual understanding in scientific applications, such as detecting st… | 29 | 1334 | active |
| ldqk/ImageSearch A .NET 10 desktop demo application that performs reverse image search (search by image) over local hard drives with tens of millions of ima… | 93 | 1329 | active |
| Gourieff/ComfyUI-ReActor ComfyUI-ReActor is a fast and simple face swap extension node for ComfyUI, based on the ReActor face-swapping engine. It includes a nudity … | 65 | 1327 | active |
| cvzone/cvzone CVZone is a Python computer vision helper library that wraps OpenCV and MediaPipe to simplify image processing and AI functions like hand t… | 32 | 1325 | active |
| neozhaoliang/surround-view-system-introduction A Python implementation of a vehicle surround-view (bird's-eye view) camera system, covering fisheye camera calibration, projection, image … | 66 | 1321 | active |
| HKUST-Aerial-Robotics/VINS-Fusion VINS-Fusion is an optimization-based multi-sensor state estimator for accurate self-localization in autonomous applications such as drones,… | 32 | 4684 | maintenance |
| hku-mars/Point-LIO Point-LIO is a robust high-bandwidth LiDAR-inertial odometry framework that estimates ego-motion and builds maps by fusing LiDAR point clou… | 71 | 1318 | active |
| open-edge-platform/geti Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru… | 98 | 1317 | active |
| seetaface/SeetaFaceEngine SeetaFace Engine is an open-source C++ face recognition engine comprising face detection, face alignment, and face identification modules. … | 32 | 4636 | maintenance |
| Vincentqyw/image-matching-webui A Gradio-based web UI that matches keypoints between two images using many state-of-the-art image matching algorithms (LoFTR, SuperGlue, Li… | 91 | 1302 | active |
| vye16/shape-of-motion Shape of Motion is a Python research codebase for 4D reconstruction of dynamic scenes from a single monocular video, based on the ICCV 2025… | 30 | 1302 | active |
| OpenTeleVision/TeleVision Open-TeleVision is an open-source immersive robot teleoperation system that streams stereoscopic visual feedback to VR headsets (Apple Visi… | 24 | 1301 | active |
| STVIR/pysot PySOT is a Python research platform by SenseTime for single object visual tracking, implementing algorithms such as SiamRPN, SiamRPN++, DaS… | 45 | 4600 | maintenance |
| PyImageSearch/imutils A Python library of convenience functions that simplify common OpenCV image processing tasks such as translation, rotation, resizing, skele… | 32 | 4590 | maintenance |
| fundamentalvision/BEVFormer BEVFormer is the official PyTorch implementation of an ECCV 2022 paper that learns bird's-eye-view (BEV) representations from multi-camera … | 23 | 4579 | maintenance |
| Parskatt/RoMa RoMa (romatch) is a Python library for robust dense feature matching between image pairs, estimating pixel-dense warps and reliable certain… | 52 | 1293 | active |
| streamlit/demo-self-driving A Streamlit demo app that provides an interactive image browser for the Udacity self-driving-car dataset with realtime YOLO object detectio… | 60 | 1290 | stable |
| BishopFox/eyeballer Eyeballer is a convolutional neural network tool that classifies screenshots of web hosts taken during large-scope penetration tests. It la… | 55 | 1290 | active |
| RoyalVane/CLAN Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation… | 32 | 1289 | stable |
| hku-mars/livox_camera_calib A C++/ROS tool from HKU MARS for automatic extrinsic calibration between high-resolution LiDAR (e.g., Livox) and cameras in targetless envi… | 32 | 1289 | stable |
| Tianxiaomo/pytorch-YOLOv4 A minimal PyTorch implementation of YOLOv4 (and YOLOv4-tiny) supporting inference and training, with tools to convert Darknet weights to Py… | 32 | 4521 | maintenance |
| zju3dv/MatchAnything MatchAnything is a deep learning model for universal cross-modality image matching, released as research code accompanying a TPAMI 2026 pap… | 64 | 1279 | active |
| AaronJackson/vrn Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses… | 32 | 4517 | maintenance |
| plemeri/transparent-background A Python tool and CLI that removes backgrounds from images and videos using the InSPyReNet deep learning model (ACCV 2022). It supports ima… | 63 | 1278 | active |
| nvidia-isaac/nvblox nvblox is a GPU-accelerated C++/Python library for real-time 3D reconstruction using TSDF and ESDF volumetric mapping, designed for robots … | 85 | 1276 | active |
| Fugtemypt123/VIGA VIGA is an analysis-by-synthesis code agent that reconstructs 3D scenes and slide layouts from images by generating and executing Blender P… | 55 | 1275 | active |
| flutter-ml/google_ml_kit_flutter A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa… | 76 | 1274 | active |
| nianticlabs/monodepth2 Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra… | 32 | 4497 | maintenance |
| Renumics/spotlight Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe… | 93 | 1272 | active |
| leoxiaobin/deep-high-resolution-net.pytorch Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai… | 32 | 4480 | maintenance |
| ToniRV/NeRF-SLAM NeRF-SLAM is a real-time dense monocular SLAM system that combines neural radiance fields (Instant-NGP) with probabilistic volumetric fusio… | 32 | 1266 | active |
| ethz-asl/rovio ROVIO (Robust Visual Inertial Odometry) is a C++ framework from ETH Zurich that estimates camera and IMU trajectory using an iterated exten… | 61 | 1262 | active |
| JonathonLuiten/TrackEval TrackEval is a Python library for evaluating multi-object tracking (MOT) algorithms, implementing metrics such as HOTA, CLEARMOT, IDF1, VAC… | 32 | 1255 | stable |
| roryclear/clearcam Clearcam is a self-hosted Python NVR that adds AI object detection, tracking, mobile notifications, and semantic search to any RTSP securit… | 86 | 1254 | active |
| ziyc/drivestudio DriveStudio is a Python framework for 3D Gaussian Splatting (3DGS) based reconstruction and simulation of dynamic urban driving scenes. It … | 40 | 1254 | active |
| FaceAISDK/FaceAISDK_Android An Android SDK for fully on-device, offline face detection, recognition, liveness detection (anti-spoofing), and 1:1, 1:N, and M:N face sea… | 98 | 1252 | active |
| Linketic/CityGaussian Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia… | 66 | 1251 | active |
| peterbraden/node-opencv Native Node.js bindings for the OpenCV computer vision library, exposing Matrices, image reading/writing, and cascades like face detection … | 23 | 4384 | maintenance |
| withoutbg/withoutbg-python A Python SDK (pip install withoutbg) for removing image backgrounds, offering a free local open-weights ONNX model and an optional paid clo… | 80 | 1246 | active |
| lpiccinelli-eth/UniDepth UniDepth is a Python library and research codebase for universal monocular metric depth estimation from single images, based on CVPR 2024 a… | 35 | 1246 | active |
| abewley/sort SORT is a barebones Python implementation of a simple online and realtime multiple object tracking algorithm for 2D video sequences, based … | 32 | 4373 | maintenance |
| ZHKKKe/MODNet MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima… | 32 | 4355 | maintenance |
| stella-cv/stella_vslam stella_vslam is a community-maintained fork of OpenVSLAM implementing a monocular, stereo, and RGBD visual SLAM system in C++. It supports … | 77 | 1238 | active |
| open-mmlab/playground OpenMMLab Playground is a central hub collecting and showcasing community projects that extend OpenMMLab libraries with Segment Anything Mo… | 30 | 1236 | active |
| xtreme1-io/xtreme1 Xtreme1 is an open-source, self-hosted data labeling and annotation platform for multimodal training data, supporting images, 3D LiDAR poin… | 62 | 1234 | active |
| stereolabs/zed-sdk The ZED SDK is a cross-platform spatial perception library for Stereolabs ZED stereo cameras, providing depth sensing, SLAM, 3D reconstruct… | 91 | 1229 | active |
| twostraws/CodeScanner CodeScanner is a SwiftUI library providing a CodeScannerView struct that scans QR codes, barcodes, and other code types using the device ca… | 64 | 1224 | active |
| MotrixLab/SMPLer-X Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an… | 59 | 1220 | stable |
| kijai/ComfyUI-segment-anything-2 A set of ComfyUI custom nodes that bring Meta's Segment Anything 2 (SAM2) models into ComfyUI workflows for promptable image and video segm… | 43 | 1214 | active |
| ifzhang/FairMOT FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification… | 32 | 4244 | maintenance |
| dexsuite/dex-retargeting A Python library of retargeting optimizers that translate human hand motion (from video or pose datasets) into robot dexterous hand joint c… | 37 | 1205 | active |
| PoseLib/PoseLib PoseLib is a C++ library of minimal solvers for calibrated camera pose estimation, covering absolute and relative pose from point and line … | 68 | 1203 | active |
| dartsim/dart DART (Dynamic Animation and Robotics Toolkit) is an open-source, research-focused C++ physics engine for robotics, animation, and machine l… | 98 | 1199 | active |
| cleanlab/cleanvision CleanVision is a Python library that automatically detects issues in image datasets, such as blurry, dark, over-exposed, or near-duplicate … | 59 | 1199 | active |
| NVlabs/alpasim AlpaSim is an open-source, Python-based autonomous vehicle simulation platform for developing and testing end-to-end AV policies in closed … | 75 | 1196 | active |
| aydinnyunus/ai-captcha-bypass A Python command-line tool that uses multimodal LLMs (GPT-4o, Gemini) to automatically solve various CAPTCHA types, including text, reCAPTC… | 61 | 1196 | active |
| zai-org/CogAgent CogAgent is an open-source vision-language model (VLM) based GUI agent that understands screen captures and natural language to automate in… | 33 | 1194 | active |
| lessthanoptimal/BoofCV BoofCV is an open-source, real-time computer vision library written entirely in Java, covering image processing, camera calibration, featur… | 86 | 1192 | active |
| Tencent-Hunyuan/HunyuanWorld-Mirror HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds… | 54 | 1191 | active |
| sair-lab/AirSLAM AirSLAM is an efficient, illumination-robust point-line visual SLAM system supporting stereo visual odometry/VIO, offline map optimization,… | 47 | 1190 | active |
| pqpo/SmartCropper An Android library for smart image cropping that automatically detects document borders using OpenCV (with an optional TensorFlow Lite HED … | 66 | 4132 | maintenance |
| msracver/Deformable-ConvNets Official MXNet implementation of Deformable Convolutional Networks (ICCV 2017) and R-FCN, including deformable convolution and ROI pooling … | 32 | 4121 | maintenance |
| VladimirYugay/Gaussian-SLAM A research implementation of a dense RGBD SLAM system that uses 3D Gaussian Splatting as its scene representation to photorealistically rec… | 27 | 1182 | active |
| ubicomplab/rPPG-Toolbox rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and… | 51 | 1178 | active |
| NVlabs/Deep_Object_Pose NVIDIA's Deep Object Pose Estimation (DOPE), a deep learning system for detecting known objects and estimating their 6-DoF pose from RGB ca… | 48 | 1178 | active |
| balancap/SSD-Tensorflow A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net… | 32 | 4101 | maintenance |
| tjiiv-cprg/EPro-PnP EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc… | 41 | 1175 | stable |
| BnanZ0/ok-nte ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated… | 78 | 1174 | active |
| chongzhou96/EdgeSAM EdgeSAM is the official PyTorch implementation of a distilled, accelerated variant of the Segment Anything Model (SAM) designed for on-devi… | 37 | 1173 | active |
| IFL-CAMP/easy_handeye A ROS package with a GUI that automates hand-eye calibration between a robot and a camera/tracking system, supporting eye-in-hand and eye-o… | 47 | 1172 | active |
| magicleap/SuperGluePretrainedNetwork SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t… | 32 | 4072 | maintenance |
| juliansteenbakker/mobile_scanner A Flutter plugin for scanning barcodes and QR codes using the device camera, backed by CameraX/ML Kit on Android, AVFoundation/Apple Vision… | 98 | 1170 | active |
| FoundationVision/GLEE GLEE is a general object foundation model for images and videos that unifies detection, segmentation, tracking, grounding, and open-world o… | 26 | 1170 | active |
| jenly1314/MLKit MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin… | 79 | 1168 | active |
| Object Detection Metrics A Python toolkit implementing the most popular metrics (AP, mAP, precision-recall curves) used to evaluate object detection algorithms, wit… | 50 | 1166 | stable |
| facebookresearch/VideoPose3D A PyTorch implementation of CVPR 2019 research on 3D human pose estimation in video using temporal convolutions over 2D keypoint trajectori… | 10 | 4052 | maintenance |
| 3D ResNets for Action Recognition A PyTorch implementation of 3D ResNet and R(2+1)D models for video action recognition, accompanying CVPR 2018 and related papers. It includ… | 23 | 4038 | maintenance |
| ucla-mobility/OpenCDA OpenCDA is an open-source Python framework for prototyping and evaluating full-stack cooperative driving automation (CDA) applications in a… | 67 | 1162 | active |
| Linaom1214/TensorRT-For-YOLO-Series A Python and C++ toolkit for running YOLO-series object detection models (YOLOv3 through YOLOv12, YOLOX) with NVIDIA TensorRT, including ON… | 44 | 1162 | active |
| MCG-NKU/E2FGVI E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It… | 32 | 1161 | stable |
| yuriy-budiyev/code-scanner An Android code scanner library built on top of ZXing that provides a customizable scanner view for reading barcodes and QR codes. It suppo… | 26 | 1160 | stable |
| DepthAnything/PromptDA Prompt Depth Anything is a Python library implementing a CVPR 2025 method for high-resolution (up to 4K) accurate metric depth estimation. … | 50 | 1159 | active |
| libuvc/libuvc libuvc is a cross-platform C library for accessing USB video devices built on top of libusb. It provides fine-grained control over UVC-comp… | 32 | 1157 | active |
| fundamentalvision/Deformable-DETR Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl… | 32 | 4015 | maintenance |
| sirius-ai/LPRNet_Pytorch A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus… | 32 | 1156 | stable |
| LazarSoft/jsqrcode A JavaScript port of the ZXing QR code scanner that decodes QR codes from images or canvas in HTML5-enabled browsers. It supports webcam-ba… | 32 | 4012 | maintenance |
| OpenGVLab/VisionLLM VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc… | 33 | 1153 | active |
| DennisLiu1993/Fastest_Image_Pattern_Matching A C++ library implementing an accelerated Normalized Cross Correlation (NCC)-based template matching and image alignment algorithm, based o… | 61 | 1151 | active |
| mangdangroboticsclub/QuadrupedRobot Mini Pupper is an open-source ROS-based quadruped robot dog kit built around Raspberry Pi, with software for SLAM, navigation, and OpenCV-b… | 67 | 1150 | active |
| naurril/SUSTechPOINTS SUSTechPOINTS is a web-based 3D point cloud annotation platform for labeling LiDAR data with 3D bounding boxes, aimed at autonomous driving… | 62 | 1150 | active |
| simpler-env/SimplerEnv SIMPLER (SimplerEnv) is a collection of simulated environments built on SAPIEN/ManiSkill for evaluating real-world robot manipulation polic… | 51 | 1147 | active |
| ShiqiYu/OpenGait OpenGait is a flexible and extensible Python framework for gait recognition research, providing implementations of state-of-the-art models … | 67 | 1146 | active |
| wh200720041/floam FLOAM is a fast and optimized Lidar Odometry And Mapping (LOAM) implementation for indoor and outdoor localization, modified from LOAM and … | 32 | 1143 | stable |
| IDEA-Research/Grounding-DINO-1.5-API Python examples and API client for Grounding DINO 1.5/1.6, IDEA Research's open-world (open-set) object detection model series hosted on De… | 25 | 1143 | active |
| MIT-SPARK/Hydra Hydra is a C++ system that incrementally builds hierarchical 3D Scene Graphs from sensor data in real time. It is developed by MIT SPARK as… | 68 | 1142 | active |
| HengyiWang/spann3r Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.… | 26 | 1141 | active |
| cvg/glue-factory Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point… | 69 | 1140 | active |
| EyeTrackVR/EyeTrackVR EyeTrackVR is a free, open-source, DIY software platform that turns affordable cameras and IR LEDs mounted inside a VR headset into an eye … | 87 | 1138 | active |