Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
CellProfiler/CellProfiler
CellProfiler is a free, open-source desktop application for quantitative analysis of biological images, letting biologists build modular im…
681135active
kzampog/cilantro
cilantro is a lean, templated C++ library for processing 3D point cloud data, offering kd-trees, normal estimation, resampling, PCA, PLY I/…
451135active
Udayraj123/OMRChecker
OMRChecker is a Python application that reads and evaluates OMR (Optical Mark Recognition) sheets scanned via a scanner or phone camera. It…
671134active
OpenGVLab/SAM-Med2D
Official implementation of SAM-Med2D, a fine-tuned Segment Anything Model (SAM) for 2D medical image segmentation, trained on the SA-Med2D-…
281134active
FlagOpen/RoboBrain2.5
RoboBrain 2.5 is an open-source embodied AI foundation model from BAAI that combines multimodal large language model capabilities with 3D s…
501132active
mgonzs13/yolo_ros
A ROS 2 wrapper for Ultralytics YOLO models (YOLOv8 through YOLO26) providing object detection, tracking, instance segmentation, human pose…
931131active
noahcao/OC_SORT
OC-SORT is a pure motion-model-based multi-object tracker for video, improving on SORT by fixing Kalman filter limitations to handle occlus…
671131stable
yohanshin/WHAM
WHAM is the official PyTorch implementation of the CVPR 2024 paper 'Reconstructing World-grounded Humans with Accurate 3D Motion'. It estim…
261130active
open-mmlab/mmtracking
MMTracking is OpenMMLab's PyTorch-based toolbox for video perception tasks, unifying video object detection, multiple object tracking, sing…
233897maintenance
brownhci/WebGazer
WebGazer.js is a JavaScript eye tracking library that uses a standard webcam to predict a user's gaze location on a web page in real time. …
653888maintenance
XieZhiFa/IdCardOCR
An Android OCR library for offline recognition of Chinese second-generation ID cards, driver's licenses, and passports. It extracts all fie…
751125active
geopavlakos/hamer
HaMeR (Hand Mesh Recovery) is a transformer-based model that reconstructs 3D hand meshes from single monocular images using the MANO parame…
561124active
xverse-engine/XScene-UEPlugin
An Unreal Engine 5 plugin for real-time visualization, management, editing, and scalable hybrid rendering of 3D Gaussian Splatting models. …
401121active
princeton-vl/RAFT-Stereo
RAFT-Stereo is a PyTorch implementation of a deep learning model for stereo matching that estimates disparity maps from stereo image pairs …
751119stable
alibaba-damo-academy/RynnVLA-002
RynnVLA-002 is a unified autoregressive Vision-Language-Action and world model that generates robot actions from text and image observation…
431119active
yangxue0827/RotationDetection
AlphaRotate is a TensorFlow-based benchmark and toolbox for rotated (oriented) object detection, implementing detectors such as R2CNN, Reti…
231118active
storyicon/comfyui_segment_anything
A ComfyUI custom node that combines GroundingDINO and Segment Anything (SAM) to segment any element in an image using semantic text prompts…
271113active
OpenKinect/libfreenect
libfreenect is a userspace driver and library for the original Microsoft Xbox Kinect sensor, providing access to RGB and depth images, moto…
233826maintenance
Anionex/agent-vision-toolkit
A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, …
791108active
princeton-vl/DPVO
DPVO is a deep learning-based visual odometry and SLAM system that estimates camera trajectories from video or image sequences using patch-…
321108active
THU-MIG/RepViT
Official PyTorch implementation of RepViT, a family of lightweight CNNs designed by integrating efficient ViT architectural designs into Mo…
191108stable
spytensor/prepare_detection_dataset
A collection of Python scripts that convert object detection datasets between common annotation formats, including CSV, LabelMe JSON, COCO,…
661107active
JEOresearch/EyeTracker
A lightweight open-source Python library for 3D eye tracking that detects and fits the pupil ellipse in eye camera video or images. It is a…
671106active
cameraui/camera.ui
camera.ui is a self-hosted, local-first video surveillance (NVR) platform for security cameras with live viewing, 24/7 recording, and on-de…
1001104active
kerberos-io/agent
Kerberos Agent is an open-source, scalable video surveillance application written in Go with a React frontend, designed to connect to IP ca…
951103active
MIT-SPARK/VGGT-SLAM
VGGT-SLAM is a dense RGB SLAM system that performs real-time feed-forward 3D scene reconstruction, optimizing on the SL(4) manifold using t…
591102active
pageauc/speed-camera
A Python3 and OpenCV application that turns a Raspberry Pi, Unix, or Windows computer with a Pi camera, USB webcam, or IP/RTSP camera into …
531101active
superxslam/SuperOdom
SuperOdometry is a lightweight C++/ROS library for LiDAR-only and LiDAR-inertial odometry and mapping, developed by CMU's AirLab. It fuses …
521101active
eszdman/PhotonCamera
PhotonCamera is an open-source Android camera app that applies enhanced computational photography image processing to captured photos. It u…
681099active
hkchengrex/Cutie
Cutie is a video object segmentation framework with object-level memory reading, a follow-up to XMem offering better consistency, robustnes…
181095active
LujiaJin/One-Pot_Multi-Frame_Denoising
Official PyTorch implementation of the One-Pot Multi-frame Denoising (OPD) method published at BMVC 2022 and extended in IJCV. It provides …
601094stable
image-js/image-js
ImageJS is a JavaScript/TypeScript library for image processing and manipulation, offering features like resizing, cropping, filtering, col…
911090stable
localai-org/depth-anything.cpp
A from-scratch C++17/ggml port of ByteDance's Depth Anything 2 and 3 models for dependency-free monocular metric depth and camera pose infe…
581090active
SimpleITK/SimpleITK
SimpleITK is a simplified C++ interface to the Insight Toolkit (ITK) for multi-dimensional image analysis, including filtering, segmentatio…
981084stable
vladmandic/face-api
FaceAPI is a JavaScript library built on TensorFlow/JS that provides AI-powered face detection, rotation tracking, face description and rec…
101083active
chaxiu/munk-ai
Munk AI (the open-source Munk Test CLI) is a local-first, self-improving AI testing engine that turns natural-language intent into product-…
801081active
auduno/headtrackr
headtrackr is a JavaScript library for real-time face tracking and head tracking via a webcam using WebRTC/getUserMedia. It estimates the u…
323701maintenance
charlesq34/pointnet2
Official TensorFlow implementation of PointNet++, a deep neural network that learns hierarchical features on 3D point clouds using metric-s…
323700maintenance
minghanqin/LangSplat
Official implementation of LangSplat, a CVPR 2024 Highlight paper that constructs a 3D language field using 3D Gaussian Splatting with CLIP…
471077active
Geekgineer/YOLOs-CPP
YOLOs-CPP is a production-ready, cross-platform C++ inference library for the YOLO model family (v5 through YOLO26), built on ONNX Runtime …
881076active
aiyaapp/AiyaEffectsAndroid
AiyaEffectsSDK is an Android demo for a face-tracking visual effects SDK that renders dynamic stickers, 3D/2D animation effects, and beauty…
371075active
lolishinshi/imsearch
A Rust-based large-scale similar image search tool that uses feature point matching (ORB features with a FAISS-style index) to find full im…
941074active
cleardusk/3DDFA
A PyTorch implementation of the TPAMI 2017 paper 'Face Alignment in Full Pose Range: A 3D Total Solution' (3DDFA). It fits a 3D Morphable M…
233677maintenance
zju3dv/InfiniDepth
InfiniDepth is a CVPR 2026 research library for monocular depth estimation that represents depth as neural implicit fields, allowing depth …
531073active
facebookresearch/CutLER
CutLER is a research codebase from Meta FAIR for training object detection and instance segmentation models without human annotations, usin…
661072active
luxonis/depthai
DepthAI is Luxonis's Python library and SDK for developing with Luxonis OAK camera hardware, enabling spatial AI and computer vision on emb…
651068active
microsoft/Biodiversity
Microsoft AI for Good Lab's biodiversity research hub providing open-source AI models and tools for wildlife monitoring and conservation, i…
881066active
gangweix/pixel-perfect-depth
Pixel-Perfect Depth is a monocular depth estimation model based on pixel-space diffusion transformers that produces flying-pixel-free depth…
491064active
rust-cv/cv
Rust CV is a mono-repo of pure-Rust computer vision crates aiming to encapsulate capabilities of OpenCV, OpenMVG, and vSLAM frameworks in c…
471063active
NVlabs/SegFormer
Official PyTorch implementation of SegFormer, a transformer-based semantic segmentation framework with a hierarchical encoder and lightweig…
323629maintenance
neka-nat/cupoch
Cupoch is a C++/Python library that implements rapid 3D data processing for robotics using CUDA, based on Open3D. It provides GPU-accelerat…
631061active
devilsen/CZXing
CZXing is a C++ port of ZXing for Android that provides WeChat-level QR code and barcode scanning, including WeChat's detection and super-r…
551060active
leggedrobotics/elevation_mapping_cupy
A GPU-accelerated elevation mapping library for robotics, built on CuPy and integrated with ROS, that fuses point clouds into multi-modal t…
891059active
YunYang1994/tensorflow-yolov3
A TensorFlow 1.x implementation of the YOLOv3 real-time object detector, reproducing the 'YOLOv3: An Incremental Improvement' paper. It sup…
233614maintenance
sb-ai-lab/EmotiEffLib
EmotiEffLib (formerly HSEmotion) is a lightweight library for facial emotion and engagement recognition in photos and videos, available in …
651057active
henry123-boy/SpaTracker
SpatialTracker is the official PyTorch implementation of a CVPR 2024 Highlight paper that tracks any 2D pixels in 3D space from RGB or RGBD…
411057active
url-kaist/patchwork-plusplus
Patchwork++ is a fast, robust, and self-adaptive ground segmentation algorithm for 3D LiDAR point clouds, published at IROS 2022. It provid…
871051active
ShenhanQian/GaussianAvatars
Official research code for GaussianAvatars, a CVPR 2024 Highlight method that creates photorealistic, fully controllable head avatars by ri…
561050active
qqlu/Entity
EntitySeg is an open-source PyTorch toolbox for open-world, high-quality image segmentation, built on Detectron2. It aggregates multiple re…
321048active
anuragxel/salt
SALT is a Python-based image labeling tool built on Meta AI's Segment Anything Model, providing a barebones GUI for annotating images with …
301048active
vectr-ucla/direct_lidar_odometry
Direct LiDAR Odometry (DLO) is a lightweight, computationally-efficient frontend LiDAR odometry package for consistent and accurate pose es…
231046stable
hku-mars/FAST-Calib
FAST-Calib is a C++ tool for fast, target-based extrinsic calibration of LiDAR-camera systems, producing accurate results in about one seco…
561044active
zju3dv/EfficientLoFTR
Efficient LoFTR is a PyTorch implementation of a semi-dense local feature matching model that matches keypoints between image pairs with sp…
401042active
foolwood/SiamMask
Official PyTorch implementation of SiamMask, a deep learning framework for fast online visual object tracking and video object segmentation…
353547maintenance
vectr-ucla/direct_lidar_inertial_odometry
DLIO is a lightweight LiDAR-inertial odometry algorithm that constructs continuous-time trajectories using a coarse-to-fine approach for pr…
531039active
lkeab/gaussian-grouping
Gaussian Grouping extends 3D Gaussian Splatting to jointly reconstruct and segment open-world 3D scenes by lifting 2D SAM masks into per-Ga…
271039stable
allenv0/AirPosture
AirPosture is an open-source iOS app that turns AirPods (and compatible Beats) with dynamic head tracking into a real-time posture coach, s…
601037active
HarborYuan/ovsam
Official PyTorch implementation of Open-Vocabulary SAM (ECCV 2024), a model that unifies SAM's interactive segmentation with CLIP's open-vo…
421033active
inclusionAI/UI-Venus
UI-Venus is a family of open-source multimodal GUI agent models (9B/27B) that perform UI element grounding and task navigation from screens…
631032active
Jumpat/SegmentAnythingin3D
SA3D is a research framework that lifts 2D Segment Anything (SAM) masks into 3D segmentation of objects within a NeRF or 3D Gaussian Splatt…
401030active
continue-revolution/sd-webui-segment-anything
A Stable Diffusion WebUI extension that integrates Segment Anything and GroundingDINO to generate segmentation masks from clicks or text pr…
303499maintenance
BlueArchiveArisHelper/BAAH
BAAH (BlueArchive Aris Helper) is an open-source Python automation script with a GUI that automatically completes daily tasks in the mobile…
901028active
DLR-RM/3DObjectTracking
A collection of C++ implementations of 3D object tracking algorithms from DLR research, including region-based 6DoF trackers (RBGT, SRT3D, …
491027active
soCzech/TransNetV2
TransNet V2 is a deep neural network for shot boundary detection in videos, achieving state-of-the-art results on benchmarks like ClipShots…
321027stable
zhyever/PatchFusion
PatchFusion is a CVPR 2024 end-to-end tile-based framework for high-resolution monocular metric depth estimation from single images. It fus…
571026active
aim-uofa/AdelaiDet
AdelaiDet is an open-source Python toolbox built on Detectron2 that implements multiple instance-level detection and recognition algorithms…
323478maintenance
koide3/small_gicp
small_gicp is a header-only C++ library with Python bindings for fast, parallelized point cloud registration algorithms including ICP, Poin…
601023active
open-mmlab/mmyolo
MMYOLO is the OpenMMLab toolbox and benchmark for the YOLO series of object detection models, implemented on PyTorch. It provides unified i…
233468maintenance
richzhang/colorization
A Python library implementing automatic colorization of grayscale photos using deep neural networks from the ECCV 2016 'Colorful Image Colo…
323461maintenance
autonomousvision/gaussian-opacity-fields
Gaussian Opacity Fields (GOF) is a Python/CUDA research implementation for efficient, adaptive surface reconstruction in unbounded scenes u…
251017active
thomwolf/Magic-Sand
Magic-Sand is a C++ openFrameworks application that operates an augmented reality sandbox by pairing a Kinect depth sensor with a projector…
231016active
maximeraafat/BlenderNeRF
BlenderNeRF is a Blender add-on that generates synthetic NeRF and Gaussian Splatting datasets with a single click, exporting renders and ca…
231014active
eragonruan/text-detection-ctpn
A TensorFlow implementation of the Connectionist Text Proposal Network (CTPN) for detecting horizontal scene text in images. It includes pr…
233429maintenance
jabcode/jabcode
JAB Code (Just Another Bar Code) is a high-capacity 2D color bar code that encodes more data than traditional black-and-white barcodes. The…
671010active
SunOner/sunone_aimbot
An AI-powered aimbot for first-person shooter games that uses YOLO object detection models (YOLOv8/v10/v12) with TensorRT/ONNX acceleration…
651010active
facebookresearch/Mask2Former
Mask2Former is the official PyTorch implementation of the CVPR 2022 paper 'Masked-attention Mask Transformer for Universal Image Segmentati…
103416maintenance
graspnet/graspnet-baseline
The official baseline deep learning model for the GraspNet-1Billion benchmark, detecting dense 6-DoF grasp poses from point clouds of clutt…
351008stable
google-research/inksight
InkSight is a Google Research system that converts photos of offline handwritten text into digital ink strokes using a ViT and mT5 encoder-…
651006active
dlbeer/quirc
Quirc is a small, dependency-free C library for extracting and decoding QR codes from images, fast enough for realtime video. It handles ro…
421004active
clovaai/CRAFT-pytorch
Official PyTorch implementation of CRAFT (Character Region Awareness for Text Detection), a scene text detector that localizes text by pred…
323398maintenance
mit-han-lab/efficientvit
A collection of efficient vision foundation models from MIT Han Lab, including EfficientViT backbones for perception, EfficientViT-SAM for …
483354maintenance
shelhamer/fcn.berkeleyvision.org
Reference implementation of Fully Convolutional Networks (FCN) for semantic segmentation from the CVPR 2015 / PAMI 2016 papers, built on Ca…
323350maintenance
tianzhi0549/FCOS
Official PyTorch implementation of FCOS, a fully convolutional one-stage, anchor-free object detector published at ICCV 2019. It provides t…
323345maintenance
hamuchiwa/AutoRCCar
An open-source project that turns a hobby RC car into an autonomous self-driving vehicle using a Raspberry Pi, Arduino, camera, and ultraso…
323344maintenance
HRNet/HRNet-Semantic-Segmentation
Official PyTorch implementation of HRNet (High-Resolution Network) and the Segmentation Transformer (OCR) approach for semantic segmentatio…
323331maintenance
pytorch-yolo-v3
A minimal PyTorch implementation of the YOLO v3 object detection algorithm, supporting detection on images and video with configurable reso…
323312maintenance
NVIDIA/flownet2-pytorch
A PyTorch implementation of FlowNet 2.0 for deep-learning-based optical flow estimation, released by NVIDIA. It provides multiple network a…
663289maintenance
anandpawara/Real_Time_Image_Animation
A real-time Python application that animates a still image (e.g., a portrait) using facial motion from a live camera or video file, built o…
323248maintenance
LBXScan
LBXScan is an iOS barcode and QR code scanning library that wraps the native AVFoundation API, ZXing, and ZBar engines behind a unified int…
323238maintenance
thearn/webcam-pulse-detector
A Python desktop application that estimates a person's heart rate in real time using only a webcam, by analyzing subtle color intensity cha…
423232maintenance

← prev page 7 / 16 next →