Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
senguptaumd/Background-Matting
Official research code for 'Background Matting: The World is Your Green Screen' (CVPR 2020), a deep network that extracts per-pixel alpha m…
324769maintenance
IrisRainbowNeko/genshin_auto_fish
A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen…
234758maintenance
liwenxi/SWIFT-AI
SWIFT-AI is a deep learning system for extremely fast gigapixel-level visual understanding in scientific applications, such as detecting st…
291334active
ldqk/ImageSearch
A .NET 10 desktop demo application that performs reverse image search (search by image) over local hard drives with tens of millions of ima…
931329active
Gourieff/ComfyUI-ReActor
ComfyUI-ReActor is a fast and simple face swap extension node for ComfyUI, based on the ReActor face-swapping engine. It includes a nudity …
651327active
cvzone/cvzone
CVZone is a Python computer vision helper library that wraps OpenCV and MediaPipe to simplify image processing and AI functions like hand t…
321325active
neozhaoliang/surround-view-system-introduction
A Python implementation of a vehicle surround-view (bird's-eye view) camera system, covering fisheye camera calibration, projection, image …
661321active
HKUST-Aerial-Robotics/VINS-Fusion
VINS-Fusion is an optimization-based multi-sensor state estimator for accurate self-localization in autonomous applications such as drones,…
324684maintenance
hku-mars/Point-LIO
Point-LIO is a robust high-bandwidth LiDAR-inertial odometry framework that estimates ego-motion and builds maps by fusing LiDAR point clou…
711318active
open-edge-platform/geti
Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru…
981317active
seetaface/SeetaFaceEngine
SeetaFace Engine is an open-source C++ face recognition engine comprising face detection, face alignment, and face identification modules. …
324636maintenance
Vincentqyw/image-matching-webui
A Gradio-based web UI that matches keypoints between two images using many state-of-the-art image matching algorithms (LoFTR, SuperGlue, Li…
911302active
vye16/shape-of-motion
Shape of Motion is a Python research codebase for 4D reconstruction of dynamic scenes from a single monocular video, based on the ICCV 2025…
301302active
OpenTeleVision/TeleVision
Open-TeleVision is an open-source immersive robot teleoperation system that streams stereoscopic visual feedback to VR headsets (Apple Visi…
241301active
STVIR/pysot
PySOT is a Python research platform by SenseTime for single object visual tracking, implementing algorithms such as SiamRPN, SiamRPN++, DaS…
454600maintenance
PyImageSearch/imutils
A Python library of convenience functions that simplify common OpenCV image processing tasks such as translation, rotation, resizing, skele…
324590maintenance
fundamentalvision/BEVFormer
BEVFormer is the official PyTorch implementation of an ECCV 2022 paper that learns bird's-eye-view (BEV) representations from multi-camera …
234579maintenance
Parskatt/RoMa
RoMa (romatch) is a Python library for robust dense feature matching between image pairs, estimating pixel-dense warps and reliable certain…
521293active
streamlit/demo-self-driving
A Streamlit demo app that provides an interactive image browser for the Udacity self-driving-car dataset with realtime YOLO object detectio…
601290stable
BishopFox/eyeballer
Eyeballer is a convolutional neural network tool that classifies screenshots of web hosts taken during large-scope penetration tests. It la…
551290active
RoyalVane/CLAN
Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation…
321289stable
hku-mars/livox_camera_calib
A C++/ROS tool from HKU MARS for automatic extrinsic calibration between high-resolution LiDAR (e.g., Livox) and cameras in targetless envi…
321289stable
Tianxiaomo/pytorch-YOLOv4
A minimal PyTorch implementation of YOLOv4 (and YOLOv4-tiny) supporting inference and training, with tools to convert Darknet weights to Py…
324521maintenance
zju3dv/MatchAnything
MatchAnything is a deep learning model for universal cross-modality image matching, released as research code accompanying a TPAMI 2026 pap…
641279active
AaronJackson/vrn
Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses…
324517maintenance
plemeri/transparent-background
A Python tool and CLI that removes backgrounds from images and videos using the InSPyReNet deep learning model (ACCV 2022). It supports ima…
631278active
nvidia-isaac/nvblox
nvblox is a GPU-accelerated C++/Python library for real-time 3D reconstruction using TSDF and ESDF volumetric mapping, designed for robots …
851276active
Fugtemypt123/VIGA
VIGA is an analysis-by-synthesis code agent that reconstructs 3D scenes and slide layouts from images by generating and executing Blender P…
551275active
flutter-ml/google_ml_kit_flutter
A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa…
761274active
nianticlabs/monodepth2
Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra…
324497maintenance
Renumics/spotlight
Renumics Spotlight is an open-source tool for interactively exploring unstructured datasets (images, audio, text, video, time-series, meshe…
931272active
leoxiaobin/deep-high-resolution-net.pytorch
Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai…
324480maintenance
ToniRV/NeRF-SLAM
NeRF-SLAM is a real-time dense monocular SLAM system that combines neural radiance fields (Instant-NGP) with probabilistic volumetric fusio…
321266active
ethz-asl/rovio
ROVIO (Robust Visual Inertial Odometry) is a C++ framework from ETH Zurich that estimates camera and IMU trajectory using an iterated exten…
611262active
JonathonLuiten/TrackEval
TrackEval is a Python library for evaluating multi-object tracking (MOT) algorithms, implementing metrics such as HOTA, CLEARMOT, IDF1, VAC…
321255stable
roryclear/clearcam
Clearcam is a self-hosted Python NVR that adds AI object detection, tracking, mobile notifications, and semantic search to any RTSP securit…
861254active
ziyc/drivestudio
DriveStudio is a Python framework for 3D Gaussian Splatting (3DGS) based reconstruction and simulation of dynamic urban driving scenes. It …
401254active
FaceAISDK/FaceAISDK_Android
An Android SDK for fully on-device, offline face detection, recognition, liveness detection (anti-spoofing), and 1:1, 1:N, and M:N face sea…
981252active
Linketic/CityGaussian
Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia…
661251active
peterbraden/node-opencv
Native Node.js bindings for the OpenCV computer vision library, exposing Matrices, image reading/writing, and cascades like face detection …
234384maintenance
withoutbg/withoutbg-python
A Python SDK (pip install withoutbg) for removing image backgrounds, offering a free local open-weights ONNX model and an optional paid clo…
801246active
lpiccinelli-eth/UniDepth
UniDepth is a Python library and research codebase for universal monocular metric depth estimation from single images, based on CVPR 2024 a…
351246active
abewley/sort
SORT is a barebones Python implementation of a simple online and realtime multiple object tracking algorithm for 2D video sequences, based …
324373maintenance
ZHKKKe/MODNet
MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima…
324355maintenance
stella-cv/stella_vslam
stella_vslam is a community-maintained fork of OpenVSLAM implementing a monocular, stereo, and RGBD visual SLAM system in C++. It supports …
771238active
open-mmlab/playground
OpenMMLab Playground is a central hub collecting and showcasing community projects that extend OpenMMLab libraries with Segment Anything Mo…
301236active
xtreme1-io/xtreme1
Xtreme1 is an open-source, self-hosted data labeling and annotation platform for multimodal training data, supporting images, 3D LiDAR poin…
621234active
stereolabs/zed-sdk
The ZED SDK is a cross-platform spatial perception library for Stereolabs ZED stereo cameras, providing depth sensing, SLAM, 3D reconstruct…
911229active
twostraws/CodeScanner
CodeScanner is a SwiftUI library providing a CodeScannerView struct that scans QR codes, barcodes, and other code types using the device ca…
641224active
MotrixLab/SMPLer-X
Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an…
591220stable
kijai/ComfyUI-segment-anything-2
A set of ComfyUI custom nodes that bring Meta's Segment Anything 2 (SAM2) models into ComfyUI workflows for promptable image and video segm…
431214active
ifzhang/FairMOT
FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification…
324244maintenance
dexsuite/dex-retargeting
A Python library of retargeting optimizers that translate human hand motion (from video or pose datasets) into robot dexterous hand joint c…
371205active
PoseLib/PoseLib
PoseLib is a C++ library of minimal solvers for calibrated camera pose estimation, covering absolute and relative pose from point and line …
681203active
dartsim/dart
DART (Dynamic Animation and Robotics Toolkit) is an open-source, research-focused C++ physics engine for robotics, animation, and machine l…
981199active
cleanlab/cleanvision
CleanVision is a Python library that automatically detects issues in image datasets, such as blurry, dark, over-exposed, or near-duplicate …
591199active
NVlabs/alpasim
AlpaSim is an open-source, Python-based autonomous vehicle simulation platform for developing and testing end-to-end AV policies in closed …
751196active
aydinnyunus/ai-captcha-bypass
A Python command-line tool that uses multimodal LLMs (GPT-4o, Gemini) to automatically solve various CAPTCHA types, including text, reCAPTC…
611196active
zai-org/CogAgent
CogAgent is an open-source vision-language model (VLM) based GUI agent that understands screen captures and natural language to automate in…
331194active
lessthanoptimal/BoofCV
BoofCV is an open-source, real-time computer vision library written entirely in Java, covering image processing, camera calibration, featur…
861192active
Tencent-Hunyuan/HunyuanWorld-Mirror
HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds…
541191active
sair-lab/AirSLAM
AirSLAM is an efficient, illumination-robust point-line visual SLAM system supporting stereo visual odometry/VIO, offline map optimization,…
471190active
pqpo/SmartCropper
An Android library for smart image cropping that automatically detects document borders using OpenCV (with an optional TensorFlow Lite HED …
664132maintenance
msracver/Deformable-ConvNets
Official MXNet implementation of Deformable Convolutional Networks (ICCV 2017) and R-FCN, including deformable convolution and ROI pooling …
324121maintenance
VladimirYugay/Gaussian-SLAM
A research implementation of a dense RGBD SLAM system that uses 3D Gaussian Splatting as its scene representation to photorealistically rec…
271182active
ubicomplab/rPPG-Toolbox
rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and…
511178active
NVlabs/Deep_Object_Pose
NVIDIA's Deep Object Pose Estimation (DOPE), a deep learning system for detecting known objects and estimating their 6-DoF pose from RGB ca…
481178active
balancap/SSD-Tensorflow
A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net…
324101maintenance
tjiiv-cprg/EPro-PnP
EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc…
411175stable
BnanZ0/ok-nte
ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated…
781174active
chongzhou96/EdgeSAM
EdgeSAM is the official PyTorch implementation of a distilled, accelerated variant of the Segment Anything Model (SAM) designed for on-devi…
371173active
IFL-CAMP/easy_handeye
A ROS package with a GUI that automates hand-eye calibration between a robot and a camera/tracking system, supporting eye-in-hand and eye-o…
471172active
magicleap/SuperGluePretrainedNetwork
SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t…
324072maintenance
juliansteenbakker/mobile_scanner
A Flutter plugin for scanning barcodes and QR codes using the device camera, backed by CameraX/ML Kit on Android, AVFoundation/Apple Vision…
981170active
FoundationVision/GLEE
GLEE is a general object foundation model for images and videos that unifies detection, segmentation, tracking, grounding, and open-world o…
261170active
jenly1314/MLKit
MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin…
791168active
Object Detection Metrics
A Python toolkit implementing the most popular metrics (AP, mAP, precision-recall curves) used to evaluate object detection algorithms, wit…
501166stable
facebookresearch/VideoPose3D
A PyTorch implementation of CVPR 2019 research on 3D human pose estimation in video using temporal convolutions over 2D keypoint trajectori…
104052maintenance
3D ResNets for Action Recognition
A PyTorch implementation of 3D ResNet and R(2+1)D models for video action recognition, accompanying CVPR 2018 and related papers. It includ…
234038maintenance
ucla-mobility/OpenCDA
OpenCDA is an open-source Python framework for prototyping and evaluating full-stack cooperative driving automation (CDA) applications in a…
671162active
Linaom1214/TensorRT-For-YOLO-Series
A Python and C++ toolkit for running YOLO-series object detection models (YOLOv3 through YOLOv12, YOLOX) with NVIDIA TensorRT, including ON…
441162active
MCG-NKU/E2FGVI
E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It…
321161stable
yuriy-budiyev/code-scanner
An Android code scanner library built on top of ZXing that provides a customizable scanner view for reading barcodes and QR codes. It suppo…
261160stable
DepthAnything/PromptDA
Prompt Depth Anything is a Python library implementing a CVPR 2025 method for high-resolution (up to 4K) accurate metric depth estimation. …
501159active
libuvc/libuvc
libuvc is a cross-platform C library for accessing USB video devices built on top of libusb. It provides fine-grained control over UVC-comp…
321157active
fundamentalvision/Deformable-DETR
Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl…
324015maintenance
sirius-ai/LPRNet_Pytorch
A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus…
321156stable
LazarSoft/jsqrcode
A JavaScript port of the ZXing QR code scanner that decodes QR codes from images or canvas in HTML5-enabled browsers. It supports webcam-ba…
324012maintenance
OpenGVLab/VisionLLM
VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc…
331153active
DennisLiu1993/Fastest_Image_Pattern_Matching
A C++ library implementing an accelerated Normalized Cross Correlation (NCC)-based template matching and image alignment algorithm, based o…
611151active
mangdangroboticsclub/QuadrupedRobot
Mini Pupper is an open-source ROS-based quadruped robot dog kit built around Raspberry Pi, with software for SLAM, navigation, and OpenCV-b…
671150active
naurril/SUSTechPOINTS
SUSTechPOINTS is a web-based 3D point cloud annotation platform for labeling LiDAR data with 3D bounding boxes, aimed at autonomous driving…
621150active
simpler-env/SimplerEnv
SIMPLER (SimplerEnv) is a collection of simulated environments built on SAPIEN/ManiSkill for evaluating real-world robot manipulation polic…
511147active
ShiqiYu/OpenGait
OpenGait is a flexible and extensible Python framework for gait recognition research, providing implementations of state-of-the-art models …
671146active
wh200720041/floam
FLOAM is a fast and optimized Lidar Odometry And Mapping (LOAM) implementation for indoor and outdoor localization, modified from LOAM and …
321143stable
IDEA-Research/Grounding-DINO-1.5-API
Python examples and API client for Grounding DINO 1.5/1.6, IDEA Research's open-world (open-set) object detection model series hosted on De…
251143active
MIT-SPARK/Hydra
Hydra is a C++ system that incrementally builds hierarchical 3D Scene Graphs from sensor data in real time. It is developed by MIT SPARK as…
681142active
HengyiWang/spann3r
Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.…
261141active
cvg/glue-factory
Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point…
691140active
EyeTrackVR/EyeTrackVR
EyeTrackVR is a free, open-source, DIY software platform that turns affordable cameras and IR LEDs mounted inside a VR headset into an eye …
871138active

← prev page 6 / 16 next →