function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| facebookresearch/perception_models Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan… | 54 | 2353 | active |
| Fate-Grand-Automata/FGA Fate/Grand Automata is a native Android app that automates battles and farming in Fate/Grand Order. It uses OpenCV for screen recognition, … | 99 | 2352 | active |
| MIT-SPARK/TEASER-plusplus TEASER++ is a fast and certifiably-robust C++ library for rigid body point cloud registration in 3D, with Python and MATLAB bindings. It es… | 48 | 2337 | stable |
| MouseLand/cellpose Cellpose is a generalist deep learning algorithm for cellular and nucleus segmentation in microscopy images, with human-in-the-loop capabil… | 86 | 2331 | active |
| NextLevel/NextLevel NextLevel is a Swift camera capture library for iOS built on AVFoundation, providing photo and video capture, multi-clip recording, ARKit i… | 71 | 2331 | active |
| gruhn/vue-qrcode-reader A set of Vue.js 3 components for detecting and decoding QR codes and other barcode formats directly in the browser. It provides QrcodeStrea… | 75 | 2307 | active |
| IDEA-Research/detrex detrex is an open-source PyTorch-based research platform and toolbox for DETR-style Transformer detection algorithms, built on top of Detec… | 41 | 2306 | active |
| PRBonn/kiss-icp KISS-ICP is a LiDAR odometry pipeline built around a simple ICP-based approach that works out of the box on most datasets without parameter… | 83 | 2303 | active |
| YvanYin/Metric3D Metric3D is the official PyTorch implementation of Metric3Dv1 and Metric3Dv2, monocular geometric foundation models that predict metric dep… | 33 | 2302 | active |
| IENT/YUView YUView is a Qt-based, cross-platform YUV video player with an advanced analytics toolset for inspecting raw video sequences. It supports ma… | 66 | 2301 | active |
| emgucv/emgucv Emgu CV is a cross-platform .NET wrapper for the OpenCV image processing library, allowing OpenCV functions to be called from .NET-compatib… | 74 | 2294 | active |
| UZ-SLAMLab/ORB_SLAM3 ORB-SLAM3 is a real-time SLAM library supporting Visual, Visual-Inertial, and Multi-Map SLAM with monocular, stereo, and RGB-D cameras usin… | 23 | 8986 | maintenance |
| andrewssobral/bgslibrary BGSLibrary is a C++ framework for background subtraction in video, offering 43 algorithms for foreground-background separation built on Ope… | 61 | 2277 | active |
| Liuziyu77/Visual-RFT Official research code for Visual-RFT and Visual-ARFT, applying GRPO-based reinforcement fine-tuning with rule-based verifiable rewards to … | 42 | 2271 | active |
| OlafenwaMoses/ImageAI ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob… | 23 | 8877 | maintenance |
| ermig1979/Simd Simd Library is a free open-source C++ image processing and machine learning library with a C API and Python wrapper. Its algorithms are ha… | 98 | 2265 | active |
| yemount/pose-animator Pose Animator is a browser-based tool that animates 2D SVG vector characters in real time using pose and face keypoints detected by PoseNet… | 32 | 8852 | maintenance |
| opendatalab/DocLayout-YOLO DocLayout-YOLO is a real-time YOLO-v10-based model for detecting document layout elements (text blocks, tables, figures, etc.) in diverse d… | 29 | 2258 | active |
| stepfun-ai/gelab-zero GELab-Zero is an open-source GUI agent framework for mobile devices, combining a 4B vision-language model (GELab-Zero-4B) with engineering … | 52 | 2257 | active |
| THU-MIG/yoloe YOLOE is the official PyTorch implementation of an open-vocabulary object detection and segmentation model presented at ICCV 2025. It unifi… | 32 | 2256 | active |
| azavea/raster-vision Raster Vision is an open source Python library and low-code framework for building computer vision models on satellite, aerial, and other l… | 61 | 2240 | active |
| Tongyi-MAI/MAI-UI Qwen-UI-Agent (MAI-UI) is a foundation GUI agent model from Alibaba's Tongyi-MAI team that unifies mobile, desktop, browser, and deep-resea… | 60 | 2228 | active |
| e2b-dev/open-computer-use An open-source AI agent that controls a secure cloud Linux desktop (via E2B Desktop Sandbox) using keyboard, mouse, and shell commands, pow… | 63 | 2224 | active |
| unrealcv/unrealcv UnrealCV is an open-source Unreal Engine plugin that connects computer vision research to virtual worlds by exposing a command API and Pyth… | 67 | 2209 | active |
| MVIG-SJTU/AlphaPose AlphaPose is an open-source real-time multi-person full-body pose estimation and tracking system built on PyTorch. It detects human keypoin… | 32 | 8596 | maintenance |
| openpnp/openpnp OpenPnP is open source software (with accompanying hardware designs) for controlling SMT pick and place machines used in PCB assembly. It c… | 81 | 2205 | active |
| spla-tam/SplaTAM SplaTAM is a dense RGB-D SLAM system that uses 3D Gaussian splatting for high-fidelity scene reconstruction and precise camera tracking fro… | 17 | 2186 | active |
| yatengLG/ISAT_with_segment_anything ISAT_with_segment_anything is an interactive semi-automatic image annotation tool built on the Segment Anything Model family (SAM, SAM2, SA… | 84 | 2166 | active |
| MRPT/mrpt MRPT is a mature C++ toolkit of libraries and applications for mobile robotics, covering SLAM, localization, probabilistic filtering, senso… | 98 | 2157 | stable |
| muskie82/MonoGS MonoGS is a dense SLAM system that applies 3D Gaussian Splatting to monocular, stereo, and RGB-D camera tracking and mapping, presented at … | 25 | 2149 | active |
| ViTAE-Transformer/ViTPose Official PyTorch implementation of ViTPose and ViTPose++, Vision Transformer models for human and generic body pose estimation from NeurIPS… | 59 | 2138 | stable |
| espressif/esp-who ESP-WHO is an image processing development platform from Espressif providing face detection, face recognition, pedestrian detection, and QR… | 67 | 2133 | active |
| yyfz/Pi3 Pi3 (π³) is a feed-forward neural network for visual geometry reconstruction that eliminates the need for a fixed reference view, using a p… | 59 | 2122 | active |
| LiheYoung/Depth-Anything Depth Anything is a monocular depth estimation foundation model trained on 1.5M labeled and 62M+ unlabeled images, released as a Python lib… | 26 | 8195 | maintenance |
| nv-tlabs/vipe ViPE is an open-source video processing engine from NVIDIA that estimates camera intrinsics, camera motion, and dense near-metric depth map… | 80 | 2092 | active |
| DepthAnything/Video-Depth-Anything Video Depth Anything is a transformer-based monocular depth estimation model for arbitrarily long videos, built on Depth Anything V2. It pr… | 41 | 2087 | active |
| 1038lab/ComfyUI-RMBG A ComfyUI custom node package for advanced image background removal and segmentation of objects, faces, clothing, and fashion elements. It … | 66 | 2086 | active |
| DanBloomberg/leptonica Leptonica is an open-source C library providing a broad set of image processing and image analysis operations, with a focus on document ima… | 74 | 2074 | stable |
| sunnypilot/sunnypilot sunnypilot is an open-source driver assistance system forked from comma.ai's openpilot, providing adaptive cruise control, automated lane c… | 98 | 2065 | active |
| marcoslucianops/DeepStream-Yolo A collection of configuration files, parsers, and conversion utilities for running YOLO-family object detection models on NVIDIA DeepStream… | 61 | 2054 | active |
| visomaster/VisoMaster VisoMaster is a Python-based desktop application for AI-powered face swapping and face editing in images and videos. It supports multiple s… | 27 | 2052 | active |
| alganzory/HaramBlur HaramBlur is a browser extension that automatically detects and blurs inappropriate images and videos on web pages using on-device machine … | 18 | 2048 | active |
| w2016561536/android_virtual_cam An Xposed module for Android that replaces the camera feed of target apps with a custom video or image. It hooks camera APIs so apps receiv… | 23 | 2047 | active |
| emilianavt/OpenSeeFace OpenSeeFace is a robust realtime face and facial landmark tracking library that runs on CPU at 30-60 fps using ONNX-converted MobileNetV3 m… | 49 | 2038 | active |
| serengil/retinaface RetinaFace is a Python library for deep learning based face detection, built on TensorFlow and derived from the insightface project's Retin… | 61 | 2027 | active |
| hgjazhgj/FGO-py A fully automatic, configuration-free, cross-platform Fate/Grand Order assistant that automates farming, event climbing, and weekly mission… | 67 | 2020 | active |
| julyx10/lap Lap is an open-source, local-first desktop photo manager for macOS, Windows, and Linux built for large personal photo libraries. It offers … | 89 | 2012 | active |
| DEIM DEIMv2 is a real-time object detection framework that extends the DEIM DETR family with DINOv3-pretrained and distilled backbones plus a Sp… | 62 | 1999 | active |
| zxing-cpp/zxing-cpp ZXing-C++ is an open-source, multi-format 1D/2D barcode image processing library written in pure C++20, ported from the Java ZXing library … | 97 | 1987 | active |
| A9T9/RPA Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible … | 96 | 1985 | active |
| xingyizhou/CenterNet CenterNet is a PyTorch implementation of the 'Objects as Points' detector, which models objects as single center points detected via keypoi… | 32 | 7573 | maintenance |
| hkchengrex/XMem XMem is a PyTorch model for semi-supervised video object segmentation that tracks objects through long videos using an Atkinson-Shiffrin-in… | 23 | 1983 | stable |
| patrikhuber/eos A lightweight, header-only 3D Morphable Face Model (3DMM) fitting library written in modern C++11/14, with Python bindings. It provides mod… | 31 | 1980 | active |
| Linzaer/Ultra-Light-Fast-Generic-Face-Detector-1MB An ultra-lightweight face detection model (~1MB FP32, ~300KB quantized) designed for edge computing devices, with slim and RFB variants tra… | 32 | 7542 | maintenance |
| KIYI671/AhabAssistantLimbusCompany AALC is a Windows desktop assistant for the game Limbus Company that automates repetitive gameplay tasks using image recognition and OCR. I… | 83 | 1971 | active |
| google-deepmind/tapnet Google DeepMind's official repository for Tracking Any Point (TAP), containing the TAP-Vid and TAPVid-3D benchmarks, the TAPIR and TAPNext … | 74 | 1968 | active |
| eriklindernoren/PyTorch-YOLOv3 A minimal PyTorch implementation of YOLOv3 supporting training, inference, and evaluation, with compatibility for YOLOv4 and YOLOv7 weights… | 32 | 7440 | maintenance |
| sicxu/Deep3DFaceRecon_pytorch A PyTorch implementation of Deep3DFaceReconstruction, a weakly-supervised CNN method for reconstructing 3D face geometry from a single imag… | 32 | 1907 | stable |
| zs1083339604/FaceWinUnlock-Tauri A Windows face-recognition unlock application built with Tauri, Vue 3, and OpenCV that injects a custom Credential Provider DLL into the Wi… | 76 | 1906 | active |
| Faceplugin-ltd/Open-Source-Face-Recognition-SDK An open-source face recognition SDK by Faceplugin providing face detection, landmark extraction, feature embedding generation, and face tem… | 64 | 1903 | active |
| MAA1999/M9A M9A is an automation assistant for the mobile game Reverse: 1999, built on MaaFramework's image-recognition and simulated-control engine. I… | 92 | 1901 | active |
| showlab/ShowUI ShowUI is an open-source, lightweight 2B vision-language-action model for GUI agents and computer use, accepted at CVPR 2025. The repositor… | 57 | 1893 | active |
| zju3dv/GVHMR GVHMR is a research codebase implementing the SIGGRAPH Asia 2024 paper 'World-Grounded Human Motion Recovery via Gravity-View Coordinates'.… | 60 | 1882 | active |
| qqwweee/keras-yolo3 A Keras (TensorFlow backend) implementation of YOLOv3 for object detection, including Darknet weight conversion, image/video detection scri… | 32 | 7114 | maintenance |
| laugh12321/TensorRT-YOLO A C++/Python deployment toolkit for running YOLO-family models (YOLOv3 through YOLO26) on NVIDIA GPUs using TensorRT, with custom plugins, … | 63 | 1880 | active |
| vt-vl-lab/3d-photo-inpainting A Python research codebase from a CVPR 2020 paper that converts a single RGB-D image into a 3D photo using layered depth inpainting. It hal… | 32 | 7093 | maintenance |
| we0091234/Chinese_license_plate_detection_recognition A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports … | 71 | 1868 | active |
| gaomingqi/Track-Anything Track-Anything is an interactive tool for video object tracking and segmentation built on Segment Anything, XMem, and E2FGVI. Users specify… | 56 | 6994 | maintenance |
| SerpentAI/SerpentAI Serpent.AI is a Python framework for building game agents—AIs and bots that learn to play any video game you own—turning games into machine… | 10 | 6992 | maintenance |
| thygate/stable-diffusion-webui-depthmap-script An extension for AUTOMATIC1111's Stable Diffusion WebUI that generates high-resolution depth maps from images using models like Marigold, M… | 32 | 1853 | active |
| ConsistentlyInconsistentYT/Pixeltovoxelprojector A Python tool that projects the motion of pixels onto a voxel representation, converting 2D pixel movement into 3D voxel space. It is a pop… | 42 | 1849 | active |
| ButzYung/SystemAnimatorOnline XR Animator is an AI-based full-body motion capture application that uses a single webcam with MediaPipe and TensorFlow.js to drive MMD/VRM… | 96 | 1837 | active |
| AlibabaResearch/AdvancedLiterateMachinery A collection of original OCR and document understanding models, algorithms, and benchmarks from Alibaba's Tongyi Lab, including models like… | 55 | 1834 | active |
| norlab-ulaval/libpointmatcher libpointmatcher is a modular C++ library implementing the Iterative Closest Point (ICP) algorithm for aligning 2D and 3D point clouds, with… | 45 | 1828 | active |
| ZFTurbo/Weighted-Boxes-Fusion A Python library implementing several methods for ensembling bounding boxes from multiple object detection models, including Non-maximum Su… | 65 | 1827 | stable |
| NVIDIA-AI-Blueprints/video-search-and-summarization NVIDIA's GPU-accelerated AI Blueprint reference architecture for building video analytics agents that search, summarize, and reason over li… | 83 | 1824 | active |
| triple-mu/YOLOv8-TensorRT A library for running YOLOv8 inference accelerated with NVIDIA TensorRT, supporting detection, segmentation, pose estimation, oriented boun… | 75 | 1804 | active |
| HuangJunJie2017/BEVDet BEVDet is a Python research codebase implementing the BEVDet series of bird's-eye-view (BEV) 3D object detection models for autonomous driv… | 23 | 1801 | active |
| allenai/ai2thor AI2-THOR is an open-source platform from the Allen Institute for AI providing near photo-realistic, interactable 3D environments (iTHOR, Ma… | 45 | 1785 | active |
| injaneity/pi-computer-use A Pi extension that gives AI agents tools to observe and operate desktop applications on macOS, Windows, and Linux. Agents can find windows… | 79 | 1777 | active |
| Totoro97/NeuS Official PyTorch implementation of NeuS, a neural implicit surface reconstruction method that learns SDF-based surfaces via volume renderin… | 32 | 1777 | stable |
| nsfw-filter/nsfw-filter A free, open-source, privacy-focused browser extension that blocks NSFW images using on-device AI classification with TensorFlow.js. It hid… | 85 | 1775 | active |
| nvidia-isaac/cuVSLAM cuVSLAM is NVIDIA's CUDA-accelerated library for real-time visual odometry and simultaneous localization and mapping (SLAM). It supports mu… | 82 | 1773 | active |
| puffinsoft/jscanify jscanify is an open-source pure JavaScript document scanning library powered by OpenCV.js. It detects and highlights documents in images an… | 73 | 1769 | active |
| ankitdhall/lidar_camera_calibration A ROS package that computes the rigid-body transformation (rotation and translation) between a LiDAR and a camera using 3D-3D point corresp… | 44 | 1766 | active |
| koide3/glim GLIM is a versatile and extensible point cloud-based 3D localization and mapping (SLAM) framework written in C++. It performs direct multi-… | 76 | 1758 | active |
| kijai/ComfyUI-Florence2 A ComfyUI custom node plugin that runs Microsoft's Florence-2 vision-language model for image captioning, object detection, segmentation, a… | 60 | 1742 | active |
| auduno/clmtrackr clmtrackr is a JavaScript library for fitting facial models to faces in videos or images using Constrained Local Models with regularized la… | 23 | 6500 | maintenance |
| SHI-Labs/OneFormer OneFormer is a universal image segmentation framework (CVPR 2023) that unifies semantic, instance, and panoptic segmentation in a single tr… | 32 | 1736 | stable |
| pupil-labs/pupil Pupil is an open source eye tracking platform consisting of applications (Pupil Capture, Player, Service) that work with Pupil Labs wearabl… | 67 | 1728 | active |
| liuruoze/EasyPR EasyPR is an open-source C++ library built on OpenCV for recognizing Chinese license plates in unconstrained situations, outputting plate c… | 23 | 6429 | maintenance |
| opendatacam/opendatacam OpenDataCam is an open-source computer vision application that detects and tracks moving objects in camera feeds or video files using YOLO/… | 58 | 1725 | active |
| MultimediaTechLab/YOLO Official MIT-licensed implementation of the YOLOv9, YOLOv7, and YOLO-RD real-time object detection models, including pre-trained weights, t… | 56 | 1723 | active |
| MaliosDark/wifi-3d-fusion WiFi-3D-Fusion is an open-source research project that estimates 3D human pose from WiFi CSI (Channel State Information) signals using deep… | 27 | 1716 | active |
| verlab/accelerated_features XFeat is a lightweight, fast learned keypoint detector and descriptor for local feature extraction and image matching, supporting both spar… | 16 | 1716 | active |
| kha-white/mokuro mokuro is a Python tool that performs text detection and OCR on Japanese manga pages and generates overlay files (.mokuro or HTML) enabling… | 86 | 1712 | active |
| whitphx/streamlit-webrtc A Python library that adds real-time video and audio streaming to Streamlit apps via WebRTC. It lets developers process live camera/microph… | 94 | 1706 | active |
| OpenAdaptAI/OpenAdapt OpenAdapt is a Python framework that compiles a demonstrated GUI workflow into an inspectable, deterministic, locally executable program fo… | 97 | 1702 | active |
| software-mansion/react-native-executorch React Native ExecuTorch is a declarative React Native library for running AI models on-device, powered by Meta's ExecuTorch runtime. It shi… | 88 | 1702 | active |
| iMoonLab/yolov13 Official PyTorch implementation of YOLOv13, a real-time object detection model family (Nano to X-Large) featuring Hypergraph-based Adaptive… | 32 | 1702 | active |