Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
facebookresearch/perception_models
Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan…
542353active
Fate-Grand-Automata/FGA
Fate/Grand Automata is a native Android app that automates battles and farming in Fate/Grand Order. It uses OpenCV for screen recognition, …
992352active
MIT-SPARK/TEASER-plusplus
TEASER++ is a fast and certifiably-robust C++ library for rigid body point cloud registration in 3D, with Python and MATLAB bindings. It es…
482337stable
MouseLand/cellpose
Cellpose is a generalist deep learning algorithm for cellular and nucleus segmentation in microscopy images, with human-in-the-loop capabil…
862331active
NextLevel/NextLevel
NextLevel is a Swift camera capture library for iOS built on AVFoundation, providing photo and video capture, multi-clip recording, ARKit i…
712331active
gruhn/vue-qrcode-reader
A set of Vue.js 3 components for detecting and decoding QR codes and other barcode formats directly in the browser. It provides QrcodeStrea…
752307active
IDEA-Research/detrex
detrex is an open-source PyTorch-based research platform and toolbox for DETR-style Transformer detection algorithms, built on top of Detec…
412306active
PRBonn/kiss-icp
KISS-ICP is a LiDAR odometry pipeline built around a simple ICP-based approach that works out of the box on most datasets without parameter…
832303active
YvanYin/Metric3D
Metric3D is the official PyTorch implementation of Metric3Dv1 and Metric3Dv2, monocular geometric foundation models that predict metric dep…
332302active
IENT/YUView
YUView is a Qt-based, cross-platform YUV video player with an advanced analytics toolset for inspecting raw video sequences. It supports ma…
662301active
emgucv/emgucv
Emgu CV is a cross-platform .NET wrapper for the OpenCV image processing library, allowing OpenCV functions to be called from .NET-compatib…
742294active
UZ-SLAMLab/ORB_SLAM3
ORB-SLAM3 is a real-time SLAM library supporting Visual, Visual-Inertial, and Multi-Map SLAM with monocular, stereo, and RGB-D cameras usin…
238986maintenance
andrewssobral/bgslibrary
BGSLibrary is a C++ framework for background subtraction in video, offering 43 algorithms for foreground-background separation built on Ope…
612277active
Liuziyu77/Visual-RFT
Official research code for Visual-RFT and Visual-ARFT, applying GRPO-based reinforcement fine-tuning with rule-based verifiable rewards to …
422271active
OlafenwaMoses/ImageAI
ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob…
238877maintenance
ermig1979/Simd
Simd Library is a free open-source C++ image processing and machine learning library with a C API and Python wrapper. Its algorithms are ha…
982265active
yemount/pose-animator
Pose Animator is a browser-based tool that animates 2D SVG vector characters in real time using pose and face keypoints detected by PoseNet…
328852maintenance
opendatalab/DocLayout-YOLO
DocLayout-YOLO is a real-time YOLO-v10-based model for detecting document layout elements (text blocks, tables, figures, etc.) in diverse d…
292258active
stepfun-ai/gelab-zero
GELab-Zero is an open-source GUI agent framework for mobile devices, combining a 4B vision-language model (GELab-Zero-4B) with engineering …
522257active
THU-MIG/yoloe
YOLOE is the official PyTorch implementation of an open-vocabulary object detection and segmentation model presented at ICCV 2025. It unifi…
322256active
azavea/raster-vision
Raster Vision is an open source Python library and low-code framework for building computer vision models on satellite, aerial, and other l…
612240active
Tongyi-MAI/MAI-UI
Qwen-UI-Agent (MAI-UI) is a foundation GUI agent model from Alibaba's Tongyi-MAI team that unifies mobile, desktop, browser, and deep-resea…
602228active
e2b-dev/open-computer-use
An open-source AI agent that controls a secure cloud Linux desktop (via E2B Desktop Sandbox) using keyboard, mouse, and shell commands, pow…
632224active
unrealcv/unrealcv
UnrealCV is an open-source Unreal Engine plugin that connects computer vision research to virtual worlds by exposing a command API and Pyth…
672209active
MVIG-SJTU/AlphaPose
AlphaPose is an open-source real-time multi-person full-body pose estimation and tracking system built on PyTorch. It detects human keypoin…
328596maintenance
openpnp/openpnp
OpenPnP is open source software (with accompanying hardware designs) for controlling SMT pick and place machines used in PCB assembly. It c…
812205active
spla-tam/SplaTAM
SplaTAM is a dense RGB-D SLAM system that uses 3D Gaussian splatting for high-fidelity scene reconstruction and precise camera tracking fro…
172186active
yatengLG/ISAT_with_segment_anything
ISAT_with_segment_anything is an interactive semi-automatic image annotation tool built on the Segment Anything Model family (SAM, SAM2, SA…
842166active
MRPT/mrpt
MRPT is a mature C++ toolkit of libraries and applications for mobile robotics, covering SLAM, localization, probabilistic filtering, senso…
982157stable
muskie82/MonoGS
MonoGS is a dense SLAM system that applies 3D Gaussian Splatting to monocular, stereo, and RGB-D camera tracking and mapping, presented at …
252149active
ViTAE-Transformer/ViTPose
Official PyTorch implementation of ViTPose and ViTPose++, Vision Transformer models for human and generic body pose estimation from NeurIPS…
592138stable
espressif/esp-who
ESP-WHO is an image processing development platform from Espressif providing face detection, face recognition, pedestrian detection, and QR…
672133active
yyfz/Pi3
Pi3 (π³) is a feed-forward neural network for visual geometry reconstruction that eliminates the need for a fixed reference view, using a p…
592122active
LiheYoung/Depth-Anything
Depth Anything is a monocular depth estimation foundation model trained on 1.5M labeled and 62M+ unlabeled images, released as a Python lib…
268195maintenance
nv-tlabs/vipe
ViPE is an open-source video processing engine from NVIDIA that estimates camera intrinsics, camera motion, and dense near-metric depth map…
802092active
DepthAnything/Video-Depth-Anything
Video Depth Anything is a transformer-based monocular depth estimation model for arbitrarily long videos, built on Depth Anything V2. It pr…
412087active
1038lab/ComfyUI-RMBG
A ComfyUI custom node package for advanced image background removal and segmentation of objects, faces, clothing, and fashion elements. It …
662086active
DanBloomberg/leptonica
Leptonica is an open-source C library providing a broad set of image processing and image analysis operations, with a focus on document ima…
742074stable
sunnypilot/sunnypilot
sunnypilot is an open-source driver assistance system forked from comma.ai's openpilot, providing adaptive cruise control, automated lane c…
982065active
marcoslucianops/DeepStream-Yolo
A collection of configuration files, parsers, and conversion utilities for running YOLO-family object detection models on NVIDIA DeepStream…
612054active
visomaster/VisoMaster
VisoMaster is a Python-based desktop application for AI-powered face swapping and face editing in images and videos. It supports multiple s…
272052active
alganzory/HaramBlur
HaramBlur is a browser extension that automatically detects and blurs inappropriate images and videos on web pages using on-device machine …
182048active
w2016561536/android_virtual_cam
An Xposed module for Android that replaces the camera feed of target apps with a custom video or image. It hooks camera APIs so apps receiv…
232047active
emilianavt/OpenSeeFace
OpenSeeFace is a robust realtime face and facial landmark tracking library that runs on CPU at 30-60 fps using ONNX-converted MobileNetV3 m…
492038active
serengil/retinaface
RetinaFace is a Python library for deep learning based face detection, built on TensorFlow and derived from the insightface project's Retin…
612027active
hgjazhgj/FGO-py
A fully automatic, configuration-free, cross-platform Fate/Grand Order assistant that automates farming, event climbing, and weekly mission…
672020active
julyx10/lap
Lap is an open-source, local-first desktop photo manager for macOS, Windows, and Linux built for large personal photo libraries. It offers …
892012active
DEIM
DEIMv2 is a real-time object detection framework that extends the DEIM DETR family with DINOv3-pretrained and distilled backbones plus a Sp…
621999active
zxing-cpp/zxing-cpp
ZXing-C++ is an open-source, multi-format 1D/2D barcode image processing library written in pure C++20, ported from the Java ZXing library …
971987active
A9T9/RPA
Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible …
961985active
xingyizhou/CenterNet
CenterNet is a PyTorch implementation of the 'Objects as Points' detector, which models objects as single center points detected via keypoi…
327573maintenance
hkchengrex/XMem
XMem is a PyTorch model for semi-supervised video object segmentation that tracks objects through long videos using an Atkinson-Shiffrin-in…
231983stable
patrikhuber/eos
A lightweight, header-only 3D Morphable Face Model (3DMM) fitting library written in modern C++11/14, with Python bindings. It provides mod…
311980active
Linzaer/Ultra-Light-Fast-Generic-Face-Detector-1MB
An ultra-lightweight face detection model (~1MB FP32, ~300KB quantized) designed for edge computing devices, with slim and RFB variants tra…
327542maintenance
KIYI671/AhabAssistantLimbusCompany
AALC is a Windows desktop assistant for the game Limbus Company that automates repetitive gameplay tasks using image recognition and OCR. I…
831971active
google-deepmind/tapnet
Google DeepMind's official repository for Tracking Any Point (TAP), containing the TAP-Vid and TAPVid-3D benchmarks, the TAPIR and TAPNext …
741968active
eriklindernoren/PyTorch-YOLOv3
A minimal PyTorch implementation of YOLOv3 supporting training, inference, and evaluation, with compatibility for YOLOv4 and YOLOv7 weights…
327440maintenance
sicxu/Deep3DFaceRecon_pytorch
A PyTorch implementation of Deep3DFaceReconstruction, a weakly-supervised CNN method for reconstructing 3D face geometry from a single imag…
321907stable
zs1083339604/FaceWinUnlock-Tauri
A Windows face-recognition unlock application built with Tauri, Vue 3, and OpenCV that injects a custom Credential Provider DLL into the Wi…
761906active
Faceplugin-ltd/Open-Source-Face-Recognition-SDK
An open-source face recognition SDK by Faceplugin providing face detection, landmark extraction, feature embedding generation, and face tem…
641903active
MAA1999/M9A
M9A is an automation assistant for the mobile game Reverse: 1999, built on MaaFramework's image-recognition and simulated-control engine. I…
921901active
showlab/ShowUI
ShowUI is an open-source, lightweight 2B vision-language-action model for GUI agents and computer use, accepted at CVPR 2025. The repositor…
571893active
zju3dv/GVHMR
GVHMR is a research codebase implementing the SIGGRAPH Asia 2024 paper 'World-Grounded Human Motion Recovery via Gravity-View Coordinates'.…
601882active
qqwweee/keras-yolo3
A Keras (TensorFlow backend) implementation of YOLOv3 for object detection, including Darknet weight conversion, image/video detection scri…
327114maintenance
laugh12321/TensorRT-YOLO
A C++/Python deployment toolkit for running YOLO-family models (YOLOv3 through YOLO26) on NVIDIA GPUs using TensorRT, with custom plugins, …
631880active
vt-vl-lab/3d-photo-inpainting
A Python research codebase from a CVPR 2020 paper that converts a single RGB-D image into a 3D photo using layered depth inpainting. It hal…
327093maintenance
we0091234/Chinese_license_plate_detection_recognition
A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports …
711868active
gaomingqi/Track-Anything
Track-Anything is an interactive tool for video object tracking and segmentation built on Segment Anything, XMem, and E2FGVI. Users specify…
566994maintenance
SerpentAI/SerpentAI
Serpent.AI is a Python framework for building game agents—AIs and bots that learn to play any video game you own—turning games into machine…
106992maintenance
thygate/stable-diffusion-webui-depthmap-script
An extension for AUTOMATIC1111's Stable Diffusion WebUI that generates high-resolution depth maps from images using models like Marigold, M…
321853active
ConsistentlyInconsistentYT/Pixeltovoxelprojector
A Python tool that projects the motion of pixels onto a voxel representation, converting 2D pixel movement into 3D voxel space. It is a pop…
421849active
ButzYung/SystemAnimatorOnline
XR Animator is an AI-based full-body motion capture application that uses a single webcam with MediaPipe and TensorFlow.js to drive MMD/VRM…
961837active
AlibabaResearch/AdvancedLiterateMachinery
A collection of original OCR and document understanding models, algorithms, and benchmarks from Alibaba's Tongyi Lab, including models like…
551834active
norlab-ulaval/libpointmatcher
libpointmatcher is a modular C++ library implementing the Iterative Closest Point (ICP) algorithm for aligning 2D and 3D point clouds, with…
451828active
ZFTurbo/Weighted-Boxes-Fusion
A Python library implementing several methods for ensembling bounding boxes from multiple object detection models, including Non-maximum Su…
651827stable
NVIDIA-AI-Blueprints/video-search-and-summarization
NVIDIA's GPU-accelerated AI Blueprint reference architecture for building video analytics agents that search, summarize, and reason over li…
831824active
triple-mu/YOLOv8-TensorRT
A library for running YOLOv8 inference accelerated with NVIDIA TensorRT, supporting detection, segmentation, pose estimation, oriented boun…
751804active
HuangJunJie2017/BEVDet
BEVDet is a Python research codebase implementing the BEVDet series of bird's-eye-view (BEV) 3D object detection models for autonomous driv…
231801active
allenai/ai2thor
AI2-THOR is an open-source platform from the Allen Institute for AI providing near photo-realistic, interactable 3D environments (iTHOR, Ma…
451785active
injaneity/pi-computer-use
A Pi extension that gives AI agents tools to observe and operate desktop applications on macOS, Windows, and Linux. Agents can find windows…
791777active
Totoro97/NeuS
Official PyTorch implementation of NeuS, a neural implicit surface reconstruction method that learns SDF-based surfaces via volume renderin…
321777stable
nsfw-filter/nsfw-filter
A free, open-source, privacy-focused browser extension that blocks NSFW images using on-device AI classification with TensorFlow.js. It hid…
851775active
nvidia-isaac/cuVSLAM
cuVSLAM is NVIDIA's CUDA-accelerated library for real-time visual odometry and simultaneous localization and mapping (SLAM). It supports mu…
821773active
puffinsoft/jscanify
jscanify is an open-source pure JavaScript document scanning library powered by OpenCV.js. It detects and highlights documents in images an…
731769active
ankitdhall/lidar_camera_calibration
A ROS package that computes the rigid-body transformation (rotation and translation) between a LiDAR and a camera using 3D-3D point corresp…
441766active
koide3/glim
GLIM is a versatile and extensible point cloud-based 3D localization and mapping (SLAM) framework written in C++. It performs direct multi-…
761758active
kijai/ComfyUI-Florence2
A ComfyUI custom node plugin that runs Microsoft's Florence-2 vision-language model for image captioning, object detection, segmentation, a…
601742active
auduno/clmtrackr
clmtrackr is a JavaScript library for fitting facial models to faces in videos or images using Constrained Local Models with regularized la…
236500maintenance
SHI-Labs/OneFormer
OneFormer is a universal image segmentation framework (CVPR 2023) that unifies semantic, instance, and panoptic segmentation in a single tr…
321736stable
pupil-labs/pupil
Pupil is an open source eye tracking platform consisting of applications (Pupil Capture, Player, Service) that work with Pupil Labs wearabl…
671728active
liuruoze/EasyPR
EasyPR is an open-source C++ library built on OpenCV for recognizing Chinese license plates in unconstrained situations, outputting plate c…
236429maintenance
opendatacam/opendatacam
OpenDataCam is an open-source computer vision application that detects and tracks moving objects in camera feeds or video files using YOLO/…
581725active
MultimediaTechLab/YOLO
Official MIT-licensed implementation of the YOLOv9, YOLOv7, and YOLO-RD real-time object detection models, including pre-trained weights, t…
561723active
MaliosDark/wifi-3d-fusion
WiFi-3D-Fusion is an open-source research project that estimates 3D human pose from WiFi CSI (Channel State Information) signals using deep…
271716active
verlab/accelerated_features
XFeat is a lightweight, fast learned keypoint detector and descriptor for local feature extraction and image matching, supporting both spar…
161716active
kha-white/mokuro
mokuro is a Python tool that performs text detection and OCR on Japanese manga pages and generates overlay files (.mokuro or HTML) enabling…
861712active
whitphx/streamlit-webrtc
A Python library that adds real-time video and audio streaming to Streamlit apps via WebRTC. It lets developers process live camera/microph…
941706active
OpenAdaptAI/OpenAdapt
OpenAdapt is a Python framework that compiles a demonstrated GUI workflow into an inspectable, deterministic, locally executable program fo…
971702active
software-mansion/react-native-executorch
React Native ExecuTorch is a declarative React Native library for running AI models on-device, powered by Meta's ExecuTorch runtime. It shi…
881702active
iMoonLab/yolov13
Official PyTorch implementation of YOLOv13, a real-time object detection model family (Nano to X-Large) featuring Hypergraph-based Adaptive…
321702active

← prev page 4 / 16 next →