function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| visionml/pytracking PyTracking is a PyTorch-based framework for visual object tracking and video object segmentation, providing official implementations of tra… | 23 | 3514 | active |
| autorope/donkeycar Donkeycar is an open-source Python library and hardware platform for building small-scale self-driving RC cars with Raspberry Pi or Jetson … | 88 | 3495 | active |
| AliceVision Meshroom is an open-source, node-based visual programming application for building and executing data processing pipelines, best known for … | 69 | 3486 | active |
| vietanhdev/anylabeling AnyLabeling is a desktop image annotation tool that combines LabelImg/Labelme-style manual labeling with AI-assisted auto-labeling. It runs… | 85 | 3463 | active |
| gkjohnson/three-mesh-bvh A Bounding Volume Hierarchy (BVH) acceleration library for three.js that speeds up raycasting and enables spatial queries against meshes. I… | 97 | 3462 | active |
| RainerKuemmerle/g2o g2o is an open-source C++ framework for optimizing graph-based nonlinear error functions, commonly used for nonlinear least squares problem… | 67 | 3461 | stable |
| facebookresearch/sam-3d-body SAM 3D Body is a promptable model for single-image full-body 3D human mesh recovery (HMR), estimating body, feet, and hand pose using the M… | 48 | 3461 | active |
| ob-f/OpenBot OpenBot is an open-source project that turns Android smartphones into the brains of low-cost robots, paired with a ~$50 electric vehicle bo… | 67 | 3442 | active |
| roflcoopter/viseron Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti… | 98 | 3431 | active |
| davidsandberg/facenet A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC… | 32 | 14343 | maintenance |
| luigifreda/pyslam pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras with a wide range of local and global fe… | 77 | 3401 | active |
| WongKinYiu/yolov7 Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques… | 23 | 14139 | maintenance |
| jenly1314/ZXingLite ZXingLite is a streamlined, fast Android library built on ZXing for scanning and generating QR codes and barcodes, with fully customizable … | 82 | 3367 | active |
| nihui/opencv-mobile opencv-mobile provides minimal, prebuilt OpenCV binary packages for Android, iOS, ARM Linux, Windows, Linux, macOS, HarmonyOS, WebAssembly,… | 85 | 3345 | active |
| Peterande/D-FINE D-FINE is the official PyTorch implementation of an ICLR 2025 Spotlight paper that redefines the regression task in DETR-style detectors as… | 67 | 3305 | active |
| XiaoMi/xiaomi-miloco Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an… | 82 | 3292 | active |
| vladmandic/human Human is a JavaScript/TypeScript library built on TensorFlow.js that combines multiple ML models for 3D face detection and recognition, bod… | 48 | 3264 | active |
| amov-lab/Prometheus Prometheus is an open-source autonomous drone software system platform built on PX4 flight controller firmware and ROS. It provides onboard… | 57 | 3239 | active |
| mit-han-lab/bevfusion BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re… | 10 | 3230 | stable |
| PJLab-ADG/SensorsCalibration OpenCalib is a C++ multi-sensor calibration toolbox for autonomous driving that calibrates IMU, LiDAR, camera, and radar sensors, both intr… | 32 | 3200 | active |
| xianfei/SysMocap SysMocap is a cross-platform, video-driven real-time motion capture system that animates 3D virtual characters from webcam footage. It rend… | 78 | 3199 | active |
| prs-eth/Marigold Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus… | 52 | 3198 | active |
| Pointcept Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan… | 76 | 3196 | active |
| ARM-software/ComputeLibrary Arm's Compute Library is a C++ collection of over 100 low-level machine learning and computer vision functions optimized for Arm Cortex-A/N… | 96 | 3183 | active |
| google-gemini/computer-use-preview A Python reference implementation of Gemini's computer-use agent that lets the model control a browser (via Playwright or Browserbase) to c… | 61 | 3181 | active |
| rmurai0610/MASt3R-SLAM MASt3R-SLAM is a real-time monocular dense SLAM system built on the MASt3R two-view 3D reconstruction prior, producing globally consistent … | 43 | 3167 | active |
| TurixAI/TuriX-CUA TuriX is an open-source computer-use agent (CUA) that lets AI models take real actions on a desktop GUI - clicking, typing, and navigating … | 70 | 3156 | active |
| cleardusk/3DDFA_V2 3DDFA_V2 is the official PyTorch implementation of the ECCV 2020 paper 'Towards Fast, Accurate and Stable 3D Dense Face Alignment'. It regr… | 23 | 3149 | stable |
| automeris-io/WebPlotDigitizer WebPlotDigitizer is a computer vision assisted web application that extracts numerical data from images of charts and plots. It has been wi… | 65 | 3146 | active |
| z-x-yang/Segment-and-Track-Anything An open-source pipeline (SAM-Track) that segments and tracks arbitrary objects in videos using the Segment Anything Model for key-frame seg… | 61 | 3134 | active |
| micjahn/ZXing.Net ZXing.Net is a .NET port of the Java ZXing barcode library that decodes and generates barcodes such as QR Code, Data Matrix, Aztec, EAN, UP… | 72 | 3088 | active |
| naver/mast3r MASt3R is the official PyTorch implementation of 'Grounding Image Matching in 3D with MASt3R' (ECCV 2024), a model that performs dense 3D r… | 37 | 3088 | active |
| zxingify/zxingify-objc ZXingObjC is a full Objective-C port of the ZXing barcode image processing library, supporting encoding and decoding of many 1D and 2D barc… | 23 | 3075 | active |
| rpng/open_vins OpenVINS is an open-source C++ platform for visual-inertial navigation research, centered on a filter-based (MSCKF/EKF) estimator that fuse… | 47 | 3057 | active |
| open-rpa/openrpa OpenRPA is a free, open-source, enterprise-grade Robotic Process Automation (RPA) tool with a visual workflow designer for Windows. It can … | 69 | 3054 | active |
| yuyuyzl/EasyVtuber EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int… | 62 | 3051 | active |
| ZQPei/deep_sort_pytorch A PyTorch implementation of the Deep SORT multi-object tracking algorithm, pairing YOLOv3/YOLOv5 (or Mask R-CNN) detectors with a CNN re-id… | 32 | 3012 | active |
| projectchrono/chrono Project Chrono is a high-performance, open-source C++ multiphysics simulation library for multibody dynamics, finite element analysis, gran… | 81 | 2993 | stable |
| iscyy/ultralyticsPro A PyTorch-based collection of improved YOLO-family object detection models (YOLOv5 through YOLOv13, RT-DETR) with pluggable modules for bac… | 48 | 2954 | active |
| wasserth/TotalSegmentator TotalSegmentator is a Python command-line tool that robustly segments over 100 anatomical structures in CT and MR images using deep learnin… | 66 | 2952 | active |
| zju3dv/LoFTR LoFTR is a detector-free local image feature matching method using Transformers, released with PyTorch inference and training code plus pre… | 32 | 2950 | stable |
| sunsmarterjie/yolov12 YOLOv12 is a PyTorch implementation of attention-centric real-time object detectors, published at NeurIPS 2025. It provides detection model… | 59 | 2947 | active |
| microsoft/table-transformer Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un… | 23 | 2939 | active |
| jeeliz/jeelizFaceFilter A lightweight JavaScript/WebGL library for real-time face detection and tracking from a camera feed via WebRTC, designed for building augme… | 46 | 2928 | active |
| langhuihui/jessibuca Jessibuca is an open-source pure HTML5 live-streaming video player built on MediaSource, WebCodecs, and WebAssembly, with WebGL rendering a… | 94 | 2893 | active |
| idootop/MagicMirror MagicMirror is a desktop application for instant AI face swapping in photos, built with Tauri. It runs entirely offline on standard hardwar… | 37 | 2891 | active |
| nimiq/qr-scanner A lightweight JavaScript/TypeScript QR code scanner library based on Cosmo Wolfe's port of Google's ZXing library. It supports webcam video… | 32 | 2886 | stable |
| NVlabs/FoundationStereo FoundationStereo is NVIDIA's official PyTorch implementation of a foundation model for zero-shot stereo depth estimation, published as a CV… | 47 | 2874 | active |
| UX-Decoder/Semantic-SAM Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired … | 33 | 2854 | active |
| openmv/openmv OpenMV is an open-source machine vision platform consisting of camera hardware firmware programmable in Python 3 (MicroPython). The firmwar… | 91 | 2850 | active |
| OpenGVLab/InternImage InternImage is a large-scale CNN-based vision foundation model that uses deformable convolutions as its core operator, released with pretra… | 28 | 2841 | stable |
| openalpr/openalpr OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete… | 23 | 11452 | maintenance |
| microsoft/MoGe MoGe is a deep learning model from Microsoft Research that recovers 3D geometry from a single open-domain image, predicting metric point ma… | 66 | 2807 | active |
| nutonomy/nuscenes-devkit The official Python devkit for the nuScenes dataset, a large-scale autonomous driving dataset from Motional. It provides dataset loading, v… | 66 | 2796 | stable |
| espressif/esp32-camera Espressif's official camera driver library for ESP32-series SoCs (ESP32, ESP32-S2, ESP32-S3), supporting a wide range of image sensors like… | 89 | 2771 | active |
| autodistill/autodistill Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab… | 29 | 2763 | active |
| infstellar/genshin_impact_assistant A multi-functional Genshin Impact auto-assist application that uses image recognition and simulated keystrokes to automate combat, domain r… | 10 | 2750 | active |
| CVCUDA/CV-CUDA CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs… | 93 | 2718 | active |
| hiukim/mind-ar-js MindAR is a web augmented reality library supporting image tracking and face tracking, written end-to-end in JavaScript with TensorFlow.js.… | 23 | 2718 | active |
| teslamotors/react-native-camera-kit A high-performance React Native camera library providing cross-platform camera capture, QR/barcode scanning, and face detection for iOS and… | 96 | 2705 | active |
| IDEA-Research/T-Rex T-Rex is the official Python API client for T-Rex2, a generic open-set object detection model that combines text and visual prompts to dete… | 48 | 2699 | active |
| torinmb/mediapipe-touchdesigner A GPU-accelerated, self-contained MediaPipe plugin for TouchDesigner that runs MediaPipe vision models (face detection, face/hand/pose trac… | 86 | 2686 | active |
| tryolabs/norfair Norfair is a lightweight, customizable Python library for real-time multi-object tracking that works with any detector outputting (x, y) co… | 31 | 2676 | stable |
| JIA-Lab-research/LISA LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati… | 31 | 2674 | active |
| princeton-vl/DROID-SLAM DROID-SLAM is a deep learning-based visual SLAM system for monocular, stereo, and RGB-D cameras that estimates camera trajectory and dense … | 41 | 2671 | active |
| SeldonIO/alibi Alibi is a Python library providing algorithms for explaining and interpreting machine learning models, including black-box, white-box, loc… | 44 | 2644 | active |
| ultralytics/yolov3 Ultralytics' PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection. It provides training, validation… | 67 | 10596 | maintenance |
| OpenStitching/stitching A Python package providing fast and robust image stitching to create panoramas, built on OpenCV's stitching module. It offers both a Python… | 88 | 2620 | active |
| antoinelame/GazeTracking A Python library that provides webcam-based eye tracking, returning pupil coordinates and gaze direction in real time using OpenCV and dlib… | 68 | 2619 | active |
| SteveMacenski/slam_toolbox Slam Toolbox is a 2D SLAM library for ROS and ROS 2 providing lifelong mapping, localization, and pose-graph manipulation for potentially m… | 93 | 2609 | active |
| luca-medeiros/lang-segment-anything A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie… | 42 | 2598 | active |
| Slicer/Slicer 3D Slicer is a free, open-source desktop platform for visualization, processing, segmentation, registration, and analysis of medical and bi… | 67 | 2595 | stable |
| Mininglamp-AI/Mano-P Mano-P is an open-source GUI-VLA (vision-language-action) agent model and SDK for edge devices, enabling purely vision-driven cross-platfor… | 55 | 2589 | active |
| BAAI-Agents/Cradle Cradle is a Python framework for General Computer Control (GCC), enabling foundation agents to perform complex computer tasks using screens… | 25 | 2573 | active |
| Tencent-Hunyuan/HY-World-2.0 HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from … | 58 | 2571 | active |
| raulmur/ORB_SLAM2 ORB-SLAM2 is a real-time SLAM library for monocular, stereo, and RGB-D cameras that computes camera trajectories and sparse 3D reconstructi… | 32 | 10223 | maintenance |
| tldev/dorso Dorso is a macOS menu bar application that monitors your posture in real time using your Mac's camera or AirPods motion sensors. When it de… | 76 | 2535 | active |
| Intent-Lab/VisionClaw VisionClaw is a real-time AI assistant app for Meta Ray-Ban smart glasses that streams camera frames and microphone audio to the Gemini Liv… | 59 | 2529 | active |
| rpautrat/SuperPoint A TensorFlow (with PyTorch conversion) implementation of the SuperPoint self-supervised interest point detector and descriptor network. It … | 41 | 2511 | stable |
| yformer/EfficientSAM EfficientSAM is an efficient image segmentation model that leverages masked image pretraining to provide a lightweight alternative to Meta'… | 27 | 2491 | active |
| sthalles/SimCLR A PyTorch reference implementation of SimCLR, a self-supervised contrastive learning framework for learning visual representations from unl… | 23 | 2491 | stable |
| ppogg/YOLOv5-Lite YOLOv5-Lite is a lightweight object detection model family evolved from YOLOv5, with models as small as ~900KB (int8) that run 10-15+ FPS o… | 23 | 2487 | active |
| ipazc/mtcnn A Python library implementing the MTCNN (Multitask Cascaded Convolutional Networks) algorithm for face detection and facial landmark alignm… | 23 | 2485 | stable |
| twistedfall/opencv-rust Rust bindings for the OpenCV computer vision library, generated automatically via Clang. It exposes OpenCV 3.4 (deprecated), 4.x, and 5.x A… | 75 | 2483 | active |
| QIN2DIM/hcaptcha-challenger A Python library that solves hCaptcha challenges using multimodal large language models and ONNX vision models (YOLO, CLIP, ResNet) integra… | 86 | 2482 | active |
| xuebinqin/U-2-Net Official PyTorch implementation of U^2-Net, a nested U-structure deep network for salient object detection, published in Pattern Recognitio… | 32 | 9853 | maintenance |
| AprilRobotics/apriltag AprilTag is a small C library implementing a visual fiducial (marker) detection system, computing the precise 3D position, orientation, and… | 71 | 2480 | stable |
| kevmo314/magic-copy Magic Copy is a browser extension (Chrome, Firefox, and Figma) that uses Meta's Segment Anything Model to segment a foreground object from … | 20 | 2458 | active |
| homuler/MediaPipeUnityPlugin A Unity native plugin that ports the MediaPipe C++ API to C#, enabling MediaPipe graphs and solutions to run inside Unity applications. It … | 63 | 2452 | active |
| hku-mars/r3live R3LIVE is a tightly-coupled LiDAR-Inertial-Visual sensor fusion framework for robust, real-time state estimation and RGB-colored 3D mapping… | 48 | 2442 | active |
| wangshub/Douyin-Bot A Python bot that automates the Douyin (TikTok China) mobile app via ADB, taking screenshots and calling a face-recognition API to auto-lik… | 32 | 9631 | maintenance |
| roboflow/inference Roboflow Inference is a Python library and self-hostable inference server for deploying computer vision models on any computer or edge devi… | 91 | 2427 | active |
| ace-trump-tech/MindPaw MindPaw is an open-source desktop quadruped robot dog built on ESP8266 with roughly ¥50 in parts, featuring voice control, gesture recognit… | 58 | 2427 | active |
| schappim/macOCR macOCR is a macOS command-line tool that captures a screen region you select and runs OCR on it, copying the recognized text (or QR/barcode… | 88 | 2426 | active |
| jyjblrd/Low-Cost-Mocap A low-cost, room-scale motion capture system built with PlayStation cameras and ESP32 hardware, used to track objects and autonomously fly … | 28 | 2421 | active |
| Smorodov/Multitarget-tracker A C++ library for multiple object tracking that combines detectors (YOLO, D-FINE, RF-DETR, MobileNet-SSD) with tracking algorithms based on… | 69 | 2415 | active |
| nv-tlabs/3dgrut NVIDIA's official implementations of 3D Gaussian Ray Tracing (3DGRT) and 3D Gaussian Unscented Transform (3DGUT), which render volumetric G… | 71 | 2390 | active |
| Cicada000/VV A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d… | 35 | 2375 | active |
| pixpark/gpupixel GPUPixel is a high-performance, cross-platform real-time image and video filter library written in C++11 and built on OpenGL/ES. It provide… | 88 | 2371 | active |
| markusfisch/BinaryEye Binary Eye is a free, open-source, ad-free barcode scanner app for Android built on the ZXing-C++ library. It reads a wide range of barcode… | 99 | 2357 | active |