Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
visionml/pytracking
PyTracking is a PyTorch-based framework for visual object tracking and video object segmentation, providing official implementations of tra…
233514active
autorope/donkeycar
Donkeycar is an open-source Python library and hardware platform for building small-scale self-driving RC cars with Raspberry Pi or Jetson …
883495active
AliceVision
Meshroom is an open-source, node-based visual programming application for building and executing data processing pipelines, best known for …
693486active
vietanhdev/anylabeling
AnyLabeling is a desktop image annotation tool that combines LabelImg/Labelme-style manual labeling with AI-assisted auto-labeling. It runs…
853463active
gkjohnson/three-mesh-bvh
A Bounding Volume Hierarchy (BVH) acceleration library for three.js that speeds up raycasting and enables spatial queries against meshes. I…
973462active
RainerKuemmerle/g2o
g2o is an open-source C++ framework for optimizing graph-based nonlinear error functions, commonly used for nonlinear least squares problem…
673461stable
facebookresearch/sam-3d-body
SAM 3D Body is a promptable model for single-image full-body 3D human mesh recovery (HMR), estimating body, feet, and hand pose using the M…
483461active
ob-f/OpenBot
OpenBot is an open-source project that turns Android smartphones into the brains of low-cost robots, paired with a ~$50 electric vehicle bo…
673442active
roflcoopter/viseron
Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti…
983431active
davidsandberg/facenet
A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC…
3214343maintenance
luigifreda/pyslam
pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras with a wide range of local and global fe…
773401active
WongKinYiu/yolov7
Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques…
2314139maintenance
jenly1314/ZXingLite
ZXingLite is a streamlined, fast Android library built on ZXing for scanning and generating QR codes and barcodes, with fully customizable …
823367active
nihui/opencv-mobile
opencv-mobile provides minimal, prebuilt OpenCV binary packages for Android, iOS, ARM Linux, Windows, Linux, macOS, HarmonyOS, WebAssembly,…
853345active
Peterande/D-FINE
D-FINE is the official PyTorch implementation of an ICLR 2025 Spotlight paper that redefines the regression task in DETR-style detectors as…
673305active
XiaoMi/xiaomi-miloco
Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an…
823292active
vladmandic/human
Human is a JavaScript/TypeScript library built on TensorFlow.js that combines multiple ML models for 3D face detection and recognition, bod…
483264active
amov-lab/Prometheus
Prometheus is an open-source autonomous drone software system platform built on PX4 flight controller firmware and ROS. It provides onboard…
573239active
mit-han-lab/bevfusion
BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re…
103230stable
PJLab-ADG/SensorsCalibration
OpenCalib is a C++ multi-sensor calibration toolbox for autonomous driving that calibrates IMU, LiDAR, camera, and radar sensors, both intr…
323200active
xianfei/SysMocap
SysMocap is a cross-platform, video-driven real-time motion capture system that animates 3D virtual characters from webcam footage. It rend…
783199active
prs-eth/Marigold
Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus…
523198active
Pointcept
Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan…
763196active
ARM-software/ComputeLibrary
Arm's Compute Library is a C++ collection of over 100 low-level machine learning and computer vision functions optimized for Arm Cortex-A/N…
963183active
google-gemini/computer-use-preview
A Python reference implementation of Gemini's computer-use agent that lets the model control a browser (via Playwright or Browserbase) to c…
613181active
rmurai0610/MASt3R-SLAM
MASt3R-SLAM is a real-time monocular dense SLAM system built on the MASt3R two-view 3D reconstruction prior, producing globally consistent …
433167active
TurixAI/TuriX-CUA
TuriX is an open-source computer-use agent (CUA) that lets AI models take real actions on a desktop GUI - clicking, typing, and navigating …
703156active
cleardusk/3DDFA_V2
3DDFA_V2 is the official PyTorch implementation of the ECCV 2020 paper 'Towards Fast, Accurate and Stable 3D Dense Face Alignment'. It regr…
233149stable
automeris-io/WebPlotDigitizer
WebPlotDigitizer is a computer vision assisted web application that extracts numerical data from images of charts and plots. It has been wi…
653146active
z-x-yang/Segment-and-Track-Anything
An open-source pipeline (SAM-Track) that segments and tracks arbitrary objects in videos using the Segment Anything Model for key-frame seg…
613134active
micjahn/ZXing.Net
ZXing.Net is a .NET port of the Java ZXing barcode library that decodes and generates barcodes such as QR Code, Data Matrix, Aztec, EAN, UP…
723088active
naver/mast3r
MASt3R is the official PyTorch implementation of 'Grounding Image Matching in 3D with MASt3R' (ECCV 2024), a model that performs dense 3D r…
373088active
zxingify/zxingify-objc
ZXingObjC is a full Objective-C port of the ZXing barcode image processing library, supporting encoding and decoding of many 1D and 2D barc…
233075active
rpng/open_vins
OpenVINS is an open-source C++ platform for visual-inertial navigation research, centered on a filter-based (MSCKF/EKF) estimator that fuse…
473057active
open-rpa/openrpa
OpenRPA is a free, open-source, enterprise-grade Robotic Process Automation (RPA) tool with a visual workflow designer for Windows. It can …
693054active
yuyuyzl/EasyVtuber
EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int…
623051active
ZQPei/deep_sort_pytorch
A PyTorch implementation of the Deep SORT multi-object tracking algorithm, pairing YOLOv3/YOLOv5 (or Mask R-CNN) detectors with a CNN re-id…
323012active
projectchrono/chrono
Project Chrono is a high-performance, open-source C++ multiphysics simulation library for multibody dynamics, finite element analysis, gran…
812993stable
iscyy/ultralyticsPro
A PyTorch-based collection of improved YOLO-family object detection models (YOLOv5 through YOLOv13, RT-DETR) with pluggable modules for bac…
482954active
wasserth/TotalSegmentator
TotalSegmentator is a Python command-line tool that robustly segments over 100 anatomical structures in CT and MR images using deep learnin…
662952active
zju3dv/LoFTR
LoFTR is a detector-free local image feature matching method using Transformers, released with PyTorch inference and training code plus pre…
322950stable
sunsmarterjie/yolov12
YOLOv12 is a PyTorch implementation of attention-centric real-time object detectors, published at NeurIPS 2025. It provides detection model…
592947active
microsoft/table-transformer
Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un…
232939active
jeeliz/jeelizFaceFilter
A lightweight JavaScript/WebGL library for real-time face detection and tracking from a camera feed via WebRTC, designed for building augme…
462928active
langhuihui/jessibuca
Jessibuca is an open-source pure HTML5 live-streaming video player built on MediaSource, WebCodecs, and WebAssembly, with WebGL rendering a…
942893active
idootop/MagicMirror
MagicMirror is a desktop application for instant AI face swapping in photos, built with Tauri. It runs entirely offline on standard hardwar…
372891active
nimiq/qr-scanner
A lightweight JavaScript/TypeScript QR code scanner library based on Cosmo Wolfe's port of Google's ZXing library. It supports webcam video…
322886stable
NVlabs/FoundationStereo
FoundationStereo is NVIDIA's official PyTorch implementation of a foundation model for zero-shot stereo depth estimation, published as a CV…
472874active
UX-Decoder/Semantic-SAM
Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired …
332854active
openmv/openmv
OpenMV is an open-source machine vision platform consisting of camera hardware firmware programmable in Python 3 (MicroPython). The firmwar…
912850active
OpenGVLab/InternImage
InternImage is a large-scale CNN-based vision foundation model that uses deformable convolutions as its core operator, released with pretra…
282841stable
openalpr/openalpr
OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete…
2311452maintenance
microsoft/MoGe
MoGe is a deep learning model from Microsoft Research that recovers 3D geometry from a single open-domain image, predicting metric point ma…
662807active
nutonomy/nuscenes-devkit
The official Python devkit for the nuScenes dataset, a large-scale autonomous driving dataset from Motional. It provides dataset loading, v…
662796stable
espressif/esp32-camera
Espressif's official camera driver library for ESP32-series SoCs (ESP32, ESP32-S2, ESP32-S3), supporting a wide range of image sensors like…
892771active
autodistill/autodistill
Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab…
292763active
infstellar/genshin_impact_assistant
A multi-functional Genshin Impact auto-assist application that uses image recognition and simulated keystrokes to automate combat, domain r…
102750active
CVCUDA/CV-CUDA
CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs…
932718active
hiukim/mind-ar-js
MindAR is a web augmented reality library supporting image tracking and face tracking, written end-to-end in JavaScript with TensorFlow.js.…
232718active
teslamotors/react-native-camera-kit
A high-performance React Native camera library providing cross-platform camera capture, QR/barcode scanning, and face detection for iOS and…
962705active
IDEA-Research/T-Rex
T-Rex is the official Python API client for T-Rex2, a generic open-set object detection model that combines text and visual prompts to dete…
482699active
torinmb/mediapipe-touchdesigner
A GPU-accelerated, self-contained MediaPipe plugin for TouchDesigner that runs MediaPipe vision models (face detection, face/hand/pose trac…
862686active
tryolabs/norfair
Norfair is a lightweight, customizable Python library for real-time multi-object tracking that works with any detector outputting (x, y) co…
312676stable
JIA-Lab-research/LISA
LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati…
312674active
princeton-vl/DROID-SLAM
DROID-SLAM is a deep learning-based visual SLAM system for monocular, stereo, and RGB-D cameras that estimates camera trajectory and dense …
412671active
SeldonIO/alibi
Alibi is a Python library providing algorithms for explaining and interpreting machine learning models, including black-box, white-box, loc…
442644active
ultralytics/yolov3
Ultralytics' PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection. It provides training, validation…
6710596maintenance
OpenStitching/stitching
A Python package providing fast and robust image stitching to create panoramas, built on OpenCV's stitching module. It offers both a Python…
882620active
antoinelame/GazeTracking
A Python library that provides webcam-based eye tracking, returning pupil coordinates and gaze direction in real time using OpenCV and dlib…
682619active
SteveMacenski/slam_toolbox
Slam Toolbox is a 2D SLAM library for ROS and ROS 2 providing lifelong mapping, localization, and pose-graph manipulation for potentially m…
932609active
luca-medeiros/lang-segment-anything
A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie…
422598active
Slicer/Slicer
3D Slicer is a free, open-source desktop platform for visualization, processing, segmentation, registration, and analysis of medical and bi…
672595stable
Mininglamp-AI/Mano-P
Mano-P is an open-source GUI-VLA (vision-language-action) agent model and SDK for edge devices, enabling purely vision-driven cross-platfor…
552589active
BAAI-Agents/Cradle
Cradle is a Python framework for General Computer Control (GCC), enabling foundation agents to perform complex computer tasks using screens…
252573active
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from …
582571active
raulmur/ORB_SLAM2
ORB-SLAM2 is a real-time SLAM library for monocular, stereo, and RGB-D cameras that computes camera trajectories and sparse 3D reconstructi…
3210223maintenance
tldev/dorso
Dorso is a macOS menu bar application that monitors your posture in real time using your Mac's camera or AirPods motion sensors. When it de…
762535active
Intent-Lab/VisionClaw
VisionClaw is a real-time AI assistant app for Meta Ray-Ban smart glasses that streams camera frames and microphone audio to the Gemini Liv…
592529active
rpautrat/SuperPoint
A TensorFlow (with PyTorch conversion) implementation of the SuperPoint self-supervised interest point detector and descriptor network. It …
412511stable
yformer/EfficientSAM
EfficientSAM is an efficient image segmentation model that leverages masked image pretraining to provide a lightweight alternative to Meta'…
272491active
sthalles/SimCLR
A PyTorch reference implementation of SimCLR, a self-supervised contrastive learning framework for learning visual representations from unl…
232491stable
ppogg/YOLOv5-Lite
YOLOv5-Lite is a lightweight object detection model family evolved from YOLOv5, with models as small as ~900KB (int8) that run 10-15+ FPS o…
232487active
ipazc/mtcnn
A Python library implementing the MTCNN (Multitask Cascaded Convolutional Networks) algorithm for face detection and facial landmark alignm…
232485stable
twistedfall/opencv-rust
Rust bindings for the OpenCV computer vision library, generated automatically via Clang. It exposes OpenCV 3.4 (deprecated), 4.x, and 5.x A…
752483active
QIN2DIM/hcaptcha-challenger
A Python library that solves hCaptcha challenges using multimodal large language models and ONNX vision models (YOLO, CLIP, ResNet) integra…
862482active
xuebinqin/U-2-Net
Official PyTorch implementation of U^2-Net, a nested U-structure deep network for salient object detection, published in Pattern Recognitio…
329853maintenance
AprilRobotics/apriltag
AprilTag is a small C library implementing a visual fiducial (marker) detection system, computing the precise 3D position, orientation, and…
712480stable
kevmo314/magic-copy
Magic Copy is a browser extension (Chrome, Firefox, and Figma) that uses Meta's Segment Anything Model to segment a foreground object from …
202458active
homuler/MediaPipeUnityPlugin
A Unity native plugin that ports the MediaPipe C++ API to C#, enabling MediaPipe graphs and solutions to run inside Unity applications. It …
632452active
hku-mars/r3live
R3LIVE is a tightly-coupled LiDAR-Inertial-Visual sensor fusion framework for robust, real-time state estimation and RGB-colored 3D mapping…
482442active
wangshub/Douyin-Bot
A Python bot that automates the Douyin (TikTok China) mobile app via ADB, taking screenshots and calling a face-recognition API to auto-lik…
329631maintenance
roboflow/inference
Roboflow Inference is a Python library and self-hostable inference server for deploying computer vision models on any computer or edge devi…
912427active
ace-trump-tech/MindPaw
MindPaw is an open-source desktop quadruped robot dog built on ESP8266 with roughly ¥50 in parts, featuring voice control, gesture recognit…
582427active
schappim/macOCR
macOCR is a macOS command-line tool that captures a screen region you select and runs OCR on it, copying the recognized text (or QR/barcode…
882426active
jyjblrd/Low-Cost-Mocap
A low-cost, room-scale motion capture system built with PlayStation cameras and ESP32 hardware, used to track objects and autonomously fly …
282421active
Smorodov/Multitarget-tracker
A C++ library for multiple object tracking that combines detectors (YOLO, D-FINE, RF-DETR, MobileNet-SSD) with tracking algorithms based on…
692415active
nv-tlabs/3dgrut
NVIDIA's official implementations of 3D Gaussian Ray Tracing (3DGRT) and 3D Gaussian Unscented Transform (3DGUT), which render volumetric G…
712390active
Cicada000/VV
A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d…
352375active
pixpark/gpupixel
GPUPixel is a high-performance, cross-platform real-time image and video filter library written in C++11 and built on OpenGL/ES. It provide…
882371active
markusfisch/BinaryEye
Binary Eye is a free, open-source, ad-free barcode scanner app for Android built on the ZXing-C++ library. It reads a wide range of barcode…
992357active

← prev page 3 / 16 next →