function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| CellProfiler/CellProfiler CellProfiler is a free, open-source desktop application for quantitative analysis of biological images, letting biologists build modular im… | 68 | 1135 | active |
| kzampog/cilantro cilantro is a lean, templated C++ library for processing 3D point cloud data, offering kd-trees, normal estimation, resampling, PCA, PLY I/… | 45 | 1135 | active |
| Udayraj123/OMRChecker OMRChecker is a Python application that reads and evaluates OMR (Optical Mark Recognition) sheets scanned via a scanner or phone camera. It… | 67 | 1134 | active |
| OpenGVLab/SAM-Med2D Official implementation of SAM-Med2D, a fine-tuned Segment Anything Model (SAM) for 2D medical image segmentation, trained on the SA-Med2D-… | 28 | 1134 | active |
| FlagOpen/RoboBrain2.5 RoboBrain 2.5 is an open-source embodied AI foundation model from BAAI that combines multimodal large language model capabilities with 3D s… | 50 | 1132 | active |
| mgonzs13/yolo_ros A ROS 2 wrapper for Ultralytics YOLO models (YOLOv8 through YOLO26) providing object detection, tracking, instance segmentation, human pose… | 93 | 1131 | active |
| noahcao/OC_SORT OC-SORT is a pure motion-model-based multi-object tracker for video, improving on SORT by fixing Kalman filter limitations to handle occlus… | 67 | 1131 | stable |
| yohanshin/WHAM WHAM is the official PyTorch implementation of the CVPR 2024 paper 'Reconstructing World-grounded Humans with Accurate 3D Motion'. It estim… | 26 | 1130 | active |
| open-mmlab/mmtracking MMTracking is OpenMMLab's PyTorch-based toolbox for video perception tasks, unifying video object detection, multiple object tracking, sing… | 23 | 3897 | maintenance |
| brownhci/WebGazer WebGazer.js is a JavaScript eye tracking library that uses a standard webcam to predict a user's gaze location on a web page in real time. … | 65 | 3888 | maintenance |
| XieZhiFa/IdCardOCR An Android OCR library for offline recognition of Chinese second-generation ID cards, driver's licenses, and passports. It extracts all fie… | 75 | 1125 | active |
| geopavlakos/hamer HaMeR (Hand Mesh Recovery) is a transformer-based model that reconstructs 3D hand meshes from single monocular images using the MANO parame… | 56 | 1124 | active |
| xverse-engine/XScene-UEPlugin An Unreal Engine 5 plugin for real-time visualization, management, editing, and scalable hybrid rendering of 3D Gaussian Splatting models. … | 40 | 1121 | active |
| princeton-vl/RAFT-Stereo RAFT-Stereo is a PyTorch implementation of a deep learning model for stereo matching that estimates disparity maps from stereo image pairs … | 75 | 1119 | stable |
| alibaba-damo-academy/RynnVLA-002 RynnVLA-002 is a unified autoregressive Vision-Language-Action and world model that generates robot actions from text and image observation… | 43 | 1119 | active |
| yangxue0827/RotationDetection AlphaRotate is a TensorFlow-based benchmark and toolbox for rotated (oriented) object detection, implementing detectors such as R2CNN, Reti… | 23 | 1118 | active |
| storyicon/comfyui_segment_anything A ComfyUI custom node that combines GroundingDINO and Segment Anything (SAM) to segment any element in an image using semantic text prompts… | 27 | 1113 | active |
| OpenKinect/libfreenect libfreenect is a userspace driver and library for the original Microsoft Xbox Kinect sensor, providing access to RGB and depth images, moto… | 23 | 3826 | maintenance |
| Anionex/agent-vision-toolkit A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, … | 79 | 1108 | active |
| princeton-vl/DPVO DPVO is a deep learning-based visual odometry and SLAM system that estimates camera trajectories from video or image sequences using patch-… | 32 | 1108 | active |
| THU-MIG/RepViT Official PyTorch implementation of RepViT, a family of lightweight CNNs designed by integrating efficient ViT architectural designs into Mo… | 19 | 1108 | stable |
| spytensor/prepare_detection_dataset A collection of Python scripts that convert object detection datasets between common annotation formats, including CSV, LabelMe JSON, COCO,… | 66 | 1107 | active |
| JEOresearch/EyeTracker A lightweight open-source Python library for 3D eye tracking that detects and fits the pupil ellipse in eye camera video or images. It is a… | 67 | 1106 | active |
| cameraui/camera.ui camera.ui is a self-hosted, local-first video surveillance (NVR) platform for security cameras with live viewing, 24/7 recording, and on-de… | 100 | 1104 | active |
| kerberos-io/agent Kerberos Agent is an open-source, scalable video surveillance application written in Go with a React frontend, designed to connect to IP ca… | 95 | 1103 | active |
| MIT-SPARK/VGGT-SLAM VGGT-SLAM is a dense RGB SLAM system that performs real-time feed-forward 3D scene reconstruction, optimizing on the SL(4) manifold using t… | 59 | 1102 | active |
| pageauc/speed-camera A Python3 and OpenCV application that turns a Raspberry Pi, Unix, or Windows computer with a Pi camera, USB webcam, or IP/RTSP camera into … | 53 | 1101 | active |
| superxslam/SuperOdom SuperOdometry is a lightweight C++/ROS library for LiDAR-only and LiDAR-inertial odometry and mapping, developed by CMU's AirLab. It fuses … | 52 | 1101 | active |
| eszdman/PhotonCamera PhotonCamera is an open-source Android camera app that applies enhanced computational photography image processing to captured photos. It u… | 68 | 1099 | active |
| hkchengrex/Cutie Cutie is a video object segmentation framework with object-level memory reading, a follow-up to XMem offering better consistency, robustnes… | 18 | 1095 | active |
| LujiaJin/One-Pot_Multi-Frame_Denoising Official PyTorch implementation of the One-Pot Multi-frame Denoising (OPD) method published at BMVC 2022 and extended in IJCV. It provides … | 60 | 1094 | stable |
| image-js/image-js ImageJS is a JavaScript/TypeScript library for image processing and manipulation, offering features like resizing, cropping, filtering, col… | 91 | 1090 | stable |
| localai-org/depth-anything.cpp A from-scratch C++17/ggml port of ByteDance's Depth Anything 2 and 3 models for dependency-free monocular metric depth and camera pose infe… | 58 | 1090 | active |
| SimpleITK/SimpleITK SimpleITK is a simplified C++ interface to the Insight Toolkit (ITK) for multi-dimensional image analysis, including filtering, segmentatio… | 98 | 1084 | stable |
| vladmandic/face-api FaceAPI is a JavaScript library built on TensorFlow/JS that provides AI-powered face detection, rotation tracking, face description and rec… | 10 | 1083 | active |
| chaxiu/munk-ai Munk AI (the open-source Munk Test CLI) is a local-first, self-improving AI testing engine that turns natural-language intent into product-… | 80 | 1081 | active |
| auduno/headtrackr headtrackr is a JavaScript library for real-time face tracking and head tracking via a webcam using WebRTC/getUserMedia. It estimates the u… | 32 | 3701 | maintenance |
| charlesq34/pointnet2 Official TensorFlow implementation of PointNet++, a deep neural network that learns hierarchical features on 3D point clouds using metric-s… | 32 | 3700 | maintenance |
| minghanqin/LangSplat Official implementation of LangSplat, a CVPR 2024 Highlight paper that constructs a 3D language field using 3D Gaussian Splatting with CLIP… | 47 | 1077 | active |
| Geekgineer/YOLOs-CPP YOLOs-CPP is a production-ready, cross-platform C++ inference library for the YOLO model family (v5 through YOLO26), built on ONNX Runtime … | 88 | 1076 | active |
| aiyaapp/AiyaEffectsAndroid AiyaEffectsSDK is an Android demo for a face-tracking visual effects SDK that renders dynamic stickers, 3D/2D animation effects, and beauty… | 37 | 1075 | active |
| lolishinshi/imsearch A Rust-based large-scale similar image search tool that uses feature point matching (ORB features with a FAISS-style index) to find full im… | 94 | 1074 | active |
| cleardusk/3DDFA A PyTorch implementation of the TPAMI 2017 paper 'Face Alignment in Full Pose Range: A 3D Total Solution' (3DDFA). It fits a 3D Morphable M… | 23 | 3677 | maintenance |
| zju3dv/InfiniDepth InfiniDepth is a CVPR 2026 research library for monocular depth estimation that represents depth as neural implicit fields, allowing depth … | 53 | 1073 | active |
| facebookresearch/CutLER CutLER is a research codebase from Meta FAIR for training object detection and instance segmentation models without human annotations, usin… | 66 | 1072 | active |
| luxonis/depthai DepthAI is Luxonis's Python library and SDK for developing with Luxonis OAK camera hardware, enabling spatial AI and computer vision on emb… | 65 | 1068 | active |
| microsoft/Biodiversity Microsoft AI for Good Lab's biodiversity research hub providing open-source AI models and tools for wildlife monitoring and conservation, i… | 88 | 1066 | active |
| gangweix/pixel-perfect-depth Pixel-Perfect Depth is a monocular depth estimation model based on pixel-space diffusion transformers that produces flying-pixel-free depth… | 49 | 1064 | active |
| rust-cv/cv Rust CV is a mono-repo of pure-Rust computer vision crates aiming to encapsulate capabilities of OpenCV, OpenMVG, and vSLAM frameworks in c… | 47 | 1063 | active |
| NVlabs/SegFormer Official PyTorch implementation of SegFormer, a transformer-based semantic segmentation framework with a hierarchical encoder and lightweig… | 32 | 3629 | maintenance |
| neka-nat/cupoch Cupoch is a C++/Python library that implements rapid 3D data processing for robotics using CUDA, based on Open3D. It provides GPU-accelerat… | 63 | 1061 | active |
| devilsen/CZXing CZXing is a C++ port of ZXing for Android that provides WeChat-level QR code and barcode scanning, including WeChat's detection and super-r… | 55 | 1060 | active |
| leggedrobotics/elevation_mapping_cupy A GPU-accelerated elevation mapping library for robotics, built on CuPy and integrated with ROS, that fuses point clouds into multi-modal t… | 89 | 1059 | active |
| YunYang1994/tensorflow-yolov3 A TensorFlow 1.x implementation of the YOLOv3 real-time object detector, reproducing the 'YOLOv3: An Incremental Improvement' paper. It sup… | 23 | 3614 | maintenance |
| sb-ai-lab/EmotiEffLib EmotiEffLib (formerly HSEmotion) is a lightweight library for facial emotion and engagement recognition in photos and videos, available in … | 65 | 1057 | active |
| henry123-boy/SpaTracker SpatialTracker is the official PyTorch implementation of a CVPR 2024 Highlight paper that tracks any 2D pixels in 3D space from RGB or RGBD… | 41 | 1057 | active |
| url-kaist/patchwork-plusplus Patchwork++ is a fast, robust, and self-adaptive ground segmentation algorithm for 3D LiDAR point clouds, published at IROS 2022. It provid… | 87 | 1051 | active |
| ShenhanQian/GaussianAvatars Official research code for GaussianAvatars, a CVPR 2024 Highlight method that creates photorealistic, fully controllable head avatars by ri… | 56 | 1050 | active |
| qqlu/Entity EntitySeg is an open-source PyTorch toolbox for open-world, high-quality image segmentation, built on Detectron2. It aggregates multiple re… | 32 | 1048 | active |
| anuragxel/salt SALT is a Python-based image labeling tool built on Meta AI's Segment Anything Model, providing a barebones GUI for annotating images with … | 30 | 1048 | active |
| vectr-ucla/direct_lidar_odometry Direct LiDAR Odometry (DLO) is a lightweight, computationally-efficient frontend LiDAR odometry package for consistent and accurate pose es… | 23 | 1046 | stable |
| hku-mars/FAST-Calib FAST-Calib is a C++ tool for fast, target-based extrinsic calibration of LiDAR-camera systems, producing accurate results in about one seco… | 56 | 1044 | active |
| zju3dv/EfficientLoFTR Efficient LoFTR is a PyTorch implementation of a semi-dense local feature matching model that matches keypoints between image pairs with sp… | 40 | 1042 | active |
| foolwood/SiamMask Official PyTorch implementation of SiamMask, a deep learning framework for fast online visual object tracking and video object segmentation… | 35 | 3547 | maintenance |
| vectr-ucla/direct_lidar_inertial_odometry DLIO is a lightweight LiDAR-inertial odometry algorithm that constructs continuous-time trajectories using a coarse-to-fine approach for pr… | 53 | 1039 | active |
| lkeab/gaussian-grouping Gaussian Grouping extends 3D Gaussian Splatting to jointly reconstruct and segment open-world 3D scenes by lifting 2D SAM masks into per-Ga… | 27 | 1039 | stable |
| allenv0/AirPosture AirPosture is an open-source iOS app that turns AirPods (and compatible Beats) with dynamic head tracking into a real-time posture coach, s… | 60 | 1037 | active |
| HarborYuan/ovsam Official PyTorch implementation of Open-Vocabulary SAM (ECCV 2024), a model that unifies SAM's interactive segmentation with CLIP's open-vo… | 42 | 1033 | active |
| inclusionAI/UI-Venus UI-Venus is a family of open-source multimodal GUI agent models (9B/27B) that perform UI element grounding and task navigation from screens… | 63 | 1032 | active |
| Jumpat/SegmentAnythingin3D SA3D is a research framework that lifts 2D Segment Anything (SAM) masks into 3D segmentation of objects within a NeRF or 3D Gaussian Splatt… | 40 | 1030 | active |
| continue-revolution/sd-webui-segment-anything A Stable Diffusion WebUI extension that integrates Segment Anything and GroundingDINO to generate segmentation masks from clicks or text pr… | 30 | 3499 | maintenance |
| BlueArchiveArisHelper/BAAH BAAH (BlueArchive Aris Helper) is an open-source Python automation script with a GUI that automatically completes daily tasks in the mobile… | 90 | 1028 | active |
| DLR-RM/3DObjectTracking A collection of C++ implementations of 3D object tracking algorithms from DLR research, including region-based 6DoF trackers (RBGT, SRT3D, … | 49 | 1027 | active |
| soCzech/TransNetV2 TransNet V2 is a deep neural network for shot boundary detection in videos, achieving state-of-the-art results on benchmarks like ClipShots… | 32 | 1027 | stable |
| zhyever/PatchFusion PatchFusion is a CVPR 2024 end-to-end tile-based framework for high-resolution monocular metric depth estimation from single images. It fus… | 57 | 1026 | active |
| aim-uofa/AdelaiDet AdelaiDet is an open-source Python toolbox built on Detectron2 that implements multiple instance-level detection and recognition algorithms… | 32 | 3478 | maintenance |
| koide3/small_gicp small_gicp is a header-only C++ library with Python bindings for fast, parallelized point cloud registration algorithms including ICP, Poin… | 60 | 1023 | active |
| open-mmlab/mmyolo MMYOLO is the OpenMMLab toolbox and benchmark for the YOLO series of object detection models, implemented on PyTorch. It provides unified i… | 23 | 3468 | maintenance |
| richzhang/colorization A Python library implementing automatic colorization of grayscale photos using deep neural networks from the ECCV 2016 'Colorful Image Colo… | 32 | 3461 | maintenance |
| autonomousvision/gaussian-opacity-fields Gaussian Opacity Fields (GOF) is a Python/CUDA research implementation for efficient, adaptive surface reconstruction in unbounded scenes u… | 25 | 1017 | active |
| thomwolf/Magic-Sand Magic-Sand is a C++ openFrameworks application that operates an augmented reality sandbox by pairing a Kinect depth sensor with a projector… | 23 | 1016 | active |
| maximeraafat/BlenderNeRF BlenderNeRF is a Blender add-on that generates synthetic NeRF and Gaussian Splatting datasets with a single click, exporting renders and ca… | 23 | 1014 | active |
| eragonruan/text-detection-ctpn A TensorFlow implementation of the Connectionist Text Proposal Network (CTPN) for detecting horizontal scene text in images. It includes pr… | 23 | 3429 | maintenance |
| jabcode/jabcode JAB Code (Just Another Bar Code) is a high-capacity 2D color bar code that encodes more data than traditional black-and-white barcodes. The… | 67 | 1010 | active |
| SunOner/sunone_aimbot An AI-powered aimbot for first-person shooter games that uses YOLO object detection models (YOLOv8/v10/v12) with TensorRT/ONNX acceleration… | 65 | 1010 | active |
| facebookresearch/Mask2Former Mask2Former is the official PyTorch implementation of the CVPR 2022 paper 'Masked-attention Mask Transformer for Universal Image Segmentati… | 10 | 3416 | maintenance |
| graspnet/graspnet-baseline The official baseline deep learning model for the GraspNet-1Billion benchmark, detecting dense 6-DoF grasp poses from point clouds of clutt… | 35 | 1008 | stable |
| google-research/inksight InkSight is a Google Research system that converts photos of offline handwritten text into digital ink strokes using a ViT and mT5 encoder-… | 65 | 1006 | active |
| dlbeer/quirc Quirc is a small, dependency-free C library for extracting and decoding QR codes from images, fast enough for realtime video. It handles ro… | 42 | 1004 | active |
| clovaai/CRAFT-pytorch Official PyTorch implementation of CRAFT (Character Region Awareness for Text Detection), a scene text detector that localizes text by pred… | 32 | 3398 | maintenance |
| mit-han-lab/efficientvit A collection of efficient vision foundation models from MIT Han Lab, including EfficientViT backbones for perception, EfficientViT-SAM for … | 48 | 3354 | maintenance |
| shelhamer/fcn.berkeleyvision.org Reference implementation of Fully Convolutional Networks (FCN) for semantic segmentation from the CVPR 2015 / PAMI 2016 papers, built on Ca… | 32 | 3350 | maintenance |
| tianzhi0549/FCOS Official PyTorch implementation of FCOS, a fully convolutional one-stage, anchor-free object detector published at ICCV 2019. It provides t… | 32 | 3345 | maintenance |
| hamuchiwa/AutoRCCar An open-source project that turns a hobby RC car into an autonomous self-driving vehicle using a Raspberry Pi, Arduino, camera, and ultraso… | 32 | 3344 | maintenance |
| HRNet/HRNet-Semantic-Segmentation Official PyTorch implementation of HRNet (High-Resolution Network) and the Segmentation Transformer (OCR) approach for semantic segmentatio… | 32 | 3331 | maintenance |
| pytorch-yolo-v3 A minimal PyTorch implementation of the YOLO v3 object detection algorithm, supporting detection on images and video with configurable reso… | 32 | 3312 | maintenance |
| NVIDIA/flownet2-pytorch A PyTorch implementation of FlowNet 2.0 for deep-learning-based optical flow estimation, released by NVIDIA. It provides multiple network a… | 66 | 3289 | maintenance |
| anandpawara/Real_Time_Image_Animation A real-time Python application that animates a still image (e.g., a portrait) using facial motion from a live camera or video file, built o… | 32 | 3248 | maintenance |
| LBXScan LBXScan is an iOS barcode and QR code scanning library that wraps the native AVFoundation API, ZXing, and ZBar engines behind a unified int… | 32 | 3238 | maintenance |
| thearn/webcam-pulse-detector A Python desktop application that estimates a person's heart rate in real time using only a webcam, by analyzing subtle color intensity cha… | 42 | 3232 | maintenance |