Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
fangfufu/Linux-Fake-Background-Webcam
A Python application that creates a virtual webcam on GNU/Linux with fake backgrounds, including background replacement, blurring, animated…
631700active
kingsic/SGQRCode
SGQRCode is an easy-to-use iOS library for scanning barcodes and QR codes, generating QR codes, and recognizing QR codes from images. It pr…
231697stable
stephansturges/WALDO
WALDO is an open-source object detection model based on a YOLOv8 backbone, trained with a synthetic data pipeline to detect people, vehicle…
321695active
mebjas/html5-qrcode
A lightweight, zero-dependency JavaScript/TypeScript library for scanning QR codes and barcodes in the browser using the device camera or l…
486213maintenance
shubham-goel/4D-Humans
4DHumans is a Python research codebase implementing HMR 2.0, a transformer-based model for 3D human mesh recovery from single images, plus …
591671active
ZrrSkywalker/Personalize-SAM
PerSAM is the official implementation of 'Personalize Segment Anything Model with One Shot', which customizes the Segment Anything Model (S…
291671active
nwojke/deep_sort
A Python implementation of Deep SORT, a multi-object tracking algorithm that extends SORT with a deep appearance descriptor for robust data…
366169maintenance
bytedance/Sa2VA
Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understa…
701666active
thtrieu/darkflow
Darkflow is a Python library that translates Darknet's YOLO neural network definitions to TensorFlow, enabling real-time object detection a…
326139maintenance
MgArcher/Text_select_captcha
A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio…
691656active
chineseocr
A Python OCR toolkit that combines YOLO3-based text detection with CRNN/Dense recognition for Chinese and English text in natural scene ima…
326123maintenance
InsightSoftwareConsortium/ITK
The Insight Toolkit (ITK) is an open-source, cross-platform C++ library with Python bindings for image analysis, providing algorithms for p…
941647stable
RQLuo/MixTeX-Latex-OCR
MixTeX is a multimodal OCR application that recognizes LaTeX formulas, tables, and mixed Chinese/English text from images, running entirely…
221637active
gaoxiang12/lightning-lm
Lightning-LM is a C++ library providing a complete 3D LiDAR SLAM system with fast LIO front-end, real-time loop closure detection, and high…
511632active
hku-mars/FAST-LIVO
FAST-LIVO is a fast, tightly-coupled sparse-direct LiDAR-Inertial-Visual Odometry system combining a LIO subsystem that registers raw point…
511631stable
HKUST-Aerial-Robotics/VINS-Mono
VINS-Mono is a real-time SLAM framework for monocular visual-inertial systems, using an optimization-based sliding window formulation for h…
326011maintenance
jeffffffli/HybrIK
HybrIK is the official PyTorch implementation of a hybrid analytical-neural inverse kinematics method for 3D human pose and shape estimatio…
231618stable
OminousIndustries/PhoneDriver
PhoneDriver is a Python-based mobile automation agent that uses Qwen3-VL vision-language models to visually understand and control Android …
381614active
Robbyant/lingbot-depth
LingBot-Depth is a PyTorch-based model and toolkit for masked depth modeling that transforms incomplete, noisy depth sensor data into metri…
561610active
meituan/YOLOv6
YOLOv6 is a single-stage object detection framework implemented in PyTorch, designed for industrial applications with a family of pretraine…
235895maintenance
Kruk2/jasna
Jasna is a GPU-accelerated tool that detects and restores mosaics in JAV videos and still images, with a native GUI, CLI, and streaming sup…
811601active
autonomousvision/transfuser
Official PyTorch implementation of TransFuser, a transformer-based multi-modal sensor fusion model for end-to-end autonomous driving, publi…
531599stable
facebookresearch/fast3r
Fast3R is the official PyTorch implementation of a CVPR 2025 model from Meta FAIR that reconstructs 3D scenes and estimates camera poses fr…
101593active
yakhyo/uniface
UniFace is a unified Python library for face analysis that bundles detection, recognition, landmark localization, face parsing, gaze estima…
881586active
Core-Mate/OpenGUI
OpenGUI is an Android GUI agent framework that lets AI agents see, understand, and operate real mobile app interfaces on physical Android d…
771583active
ai-forever/ghost
GHOST (Generative High-fidelity One Shot Transfer) is a one-shot face swap pipeline for images and videos, published as an IEEE paper and i…
261582active
Layout-Parser/layout-parser
LayoutParser is a Python toolkit for deep learning based document image analysis, offering unified APIs for layout detection models, layout…
235774maintenance
Tencent/DepthCrafter
DepthCrafter is a diffusion-based video depth estimation model from Tencent AI Lab that generates temporally consistent long depth sequence…
381574active
Babyhamsta/Aimmy
Aimmy is a universal AI-based aim alignment mechanism (aim assist) for gamers with impairments, built in C# using YOLOv8 models run via ONN…
781568active
IDEA-Research/Rex-Omni
Rex-Omni is a 3B-parameter multimodal large language model that unifies object detection, OCR, pointing, keypoint detection, and visual pro…
471561active
real-stanford/universal_manipulation_interface
Universal Manipulation Interface (UMI) is a data collection and policy learning framework that transfers in-the-wild human demonstrations i…
631560active
ORB-HD/deface
deface is a Python command-line tool that automatically anonymizes human faces in videos and photos. It detects faces in each frame and app…
231558stable
yeemachine/kalidokit
KalidoKit is a TypeScript library that converts 3D landmark outputs from Mediapipe/Tensorflow.js face, pose, and hand tracking models into …
495699maintenance
ethz-asl/kalibr
Kalibr is a visual-inertial calibration toolbox for camera systems and inertial measurement units. It supports multi-camera, camera-IMU, IM…
325680maintenance
hustvl/MapTR
MapTR is an end-to-end transformer-based framework for online vectorized HD map construction from camera imagery in autonomous driving. It …
271542active
dindin0497/SeeIt
SeeIt is an inclusive Android app with two accessibility modes: one that converts spoken speech into text and plays corresponding ASL (Amer…
411540active
Arthur151/ROMP
ROMP is a PyTorch-based library and pip-installable API (simple-romp) for real-time monocular multi-person 3D human mesh recovery, implemen…
231538stable
studyhelperhelper/studyhelper
An Android app disguised as a Sudoku game that automates earning daily points in the Xuexi Qiangguo (学习强国) app. It uses accessibility servi…
231533active
JingyunLiang/SwinIR
Official PyTorch implementation of SwinIR, a Swin Transformer-based model for image restoration tasks including super-resolution, denoising…
235580maintenance
NirAharon/BoT-SORT
BoT-SORT is a state-of-the-art multi-object tracker that combines motion and appearance information with camera motion compensation and an …
321522active
Tencent/TFace
TFace is a research platform from Tencent Youtu Lab for trusty face analysis, covering face recognition, face security (anti-spoofing), fac…
571521active
caiyuanhao1998/Retinexformer
Retinexformer is a one-stage Retinex-based Transformer model and toolbox for low-light image enhancement, published at ICCV 2023. It suppor…
661518active
NVlabs/describe-anything
Describe Anything Model (DAM) is a vision-language model that generates detailed descriptions of user-specified regions in images and video…
321514active
BrokenSource/DepthFlow
DepthFlow is a free, open-source Python application and library that converts still images into 3D parallax effect videos using monocular d…
851512active
OneDragon-Anything/StarRailOneDragon
StarRailOneDragon is a Python-based automation application for the game Honkai: Star Rail that handles daily tasks automatically, using scr…
841512active
hkchengrex/Tracking-Anything-with-DEVA
DEVA is a decoupled video segmentation framework that combines task-specific image-level segmentation models with a universal bi-directiona…
271508stable
koide3/direct_visual_lidar_calibration
A C++ toolbox for target-less, single-shot extrinsic calibration between LiDAR sensors and cameras, supporting spinning and non-repetitive …
741507stable
czczup/ViT-Adapter
Official PyTorch implementation of ViT-Adapter, an ICLR 2023 Spotlight paper introducing a pre-training-free adapter that lets plain Vision…
341503stable
zapdos-labs/unblink
Unblink is an AI-powered camera monitoring application that uses a vision language model (Qwen3-VL) to analyze camera frames, summarize act…
501501active
meiqua/shape_based_matching
A C++ library implementing Halcon-style shape-based matching (equivalent to LINE-MOD) using gradient orientation templates for robust 2D ob…
321500active
introlab/rtabmap_ros
RTAB-Map's ROS package providing real-time appearance-based SLAM (RGB-D, stereo, and LiDAR graph SLAM) as ROS 1 and ROS 2 nodes. It integra…
761498active
ANTsX/ANTs
Advanced Normalization Tools (ANTs) is a C++ command-line library for high-dimensional medical image registration and segmentation, built o…
831497active
cheind/py-motmetrics
py-motmetrics is a Python library for evaluating multiple object tracking (MOT) results with MOTChallenge-aligned CLEAR MOT, Identity, and …
651487active
CUT3R/CUT3R
CUT3R is the official PyTorch implementation of 'Continuous 3D Perception Model with Persistent State' (CVPR 2025 Oral), a stateful recurre…
381486active
damiafuentes/DJITelloPy
A Python library wrapping the official DJI Tello and Tello EDU SDKs, implementing all Tello commands including video streaming, state packe…
241482active
hustvl/DiffusionDrive
DiffusionDrive is a truncated diffusion model for real-time end-to-end autonomous driving, released as the official PyTorch implementation …
441480active
pur1fying/blue_archive_auto_script
BAAS (Blue Archive Auto Script) is a GUI-based automation program for the mobile game Blue Archive that runs against 16:9 emulator screens.…
721476active
google/GNM
GNM is an open ecosystem of parametric statistical human models and perception stacks from Google, starting with GNM Head, a high-fidelity …
591469active
autonomousvision/mip-splatting
Mip-Splatting is a research implementation of alias-free 3D Gaussian Splatting, introducing a 3D smoothing filter and 2D Mip filter to elim…
271466active
dbolya/yolact
YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++…
515241maintenance
zylo117/Yet-Another-EfficientDet-Pytorch
A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim…
235238maintenance
NVIDIA-ISAAC-ROS/isaac_ros_visual_slam
Isaac ROS Visual SLAM is a ROS 2 package providing GPU-accelerated visual simultaneous localization and mapping (VSLAM) using stereo visual…
941448active
phonowell/genshin-impact-script
A Genshin Impact automation script written in AutoHotkey that provides features like automatic fishing, item pickup, and dialogue skipping.…
631446active
amdegroot/ssd.pytorch
A PyTorch implementation of the Single Shot MultiBox Detector (SSD) object detection model from the 2016 paper by Wei Liu et al. It include…
325221maintenance
serratus/quaggaJS
QuaggaJS is a barcode-scanner library written entirely in JavaScript that supports real-time localization and decoding of barcode types suc…
235207maintenance
sdcb/PaddleSharp
A .NET/C# wrapper around Baidu's PaddleInference C API, providing PaddleOCR, PaddleDetection, rotation detection, Chinese segmentation, and…
721441active
ThoughtfulDev/EagleEye
EagleEye is a Python-based OSINT tool that identifies social media profiles (Instagram, Facebook, Twitter, YouTube) of a person using face …
325197maintenance
valentinfrlch/ha-llmvision
LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live …
901440active
Walter0807/MotionBERT
Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske…
651439active
jakowenko/double-take
Double Take is a self-hosted Docker application providing a unified UI and API for facial recognition. It abstracts multiple face detection…
401436active
NVlabs/Fast-FoundationStereo
Fast-FoundationStereo is NVIDIA's official PyTorch implementation of a real-time zero-shot stereo matching model family, accepted to CVPR 2…
541432active
qupath/qupath
QuPath is an open-source desktop application for bioimage analysis, aimed especially at digital pathology and whole-slide imaging. It provi…
791430active
jasonmayes/Real-Time-Person-Removal
A browser-based demo that removes people from complex video backgrounds in real time using TensorFlow.js. It learns the static background o…
235154maintenance
natario1/CameraView
CameraView is a well-documented, high-level Android library that simplifies capturing pictures and videos, wrapping Camera1 and Camera2 API…
235125maintenance
ZheC/Realtime_Multi-Person_Pose_Estimation
Reference implementation of the CVPR'17 paper 'Realtime Multi-Person Pose Estimation', a bottom-up approach that detects keypoints for mult…
325123maintenance
facebookresearch/vggsfm
VGGSfM is a deep learning-based Structure from Motion pipeline from Meta AI and Oxford VGG that recovers camera poses and 3D point clouds f…
291421active
CSAILVision/semantic-segmentation-pytorch
A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra…
325078maintenance
khanamiryan/php-qrcode-detector-decoder
A pure PHP library for detecting and decoding QR codes from images, ported from the ZXing library. It works without any PHP extensions beyo…
371412active
NVlabs/BundleSDF
BundleSDF is a CVPR 2023 research implementation from NVIDIA for near real-time 6-DoF pose tracking of unknown rigid objects from monocular…
651410stable
IDEA-Research/DINO-X-API
DINO-X API is a Python client library and examples for accessing DINO-X, a hosted unified vision model for open-world object detection and …
361410active
OpenBMB/AgentCPM-GUI
AgentCPM-GUI is an open-source 8B-parameter on-device GUI agent built on MiniCPM-V that takes Android screenshots as input and autonomously…
461407active
imagej/imagej2
ImageJ2 is an open-source Java framework and application for processing and analyzing N-dimensional scientific image data, built on the Img…
751399stable
nasa/astrobee
NASA's open-source flight software for the Astrobee free-flying robots operating aboard the International Space Station, written primarily …
481397active
yfeng95/PRNet
PRNet is a Python/TensorFlow implementation of the ECCV 2018 Position Map Regression Network for joint 3D face reconstruction and dense ali…
325013maintenance
om-ai-lab/OmDet
OmDet-Turbo is a PyTorch implementation of a transformer-based open-vocabulary object detection model that detects arbitrary user-defined o…
571393active
yipianfengye/android-zxingLibrary
An Android library wrapping ZXing that lets developers integrate QR code and barcode scanning into their apps with just a few lines of code…
324995maintenance
Junyi42/monst3r
MonST3R is the official PyTorch implementation of an ICLR 2025 paper that estimates per-timestep geometry (pointmaps) from dynamic videos i…
361386active
zhixuhao/unet
A Keras implementation of the U-Net convolutional network architecture for image segmentation, based on the original biomedical segmentatio…
664941maintenance
autonomousvision/unimatch
UniMatch is a PyTorch research library implementing a unified transformer-based model for optical flow, stereo matching, and depth estimati…
321379stable
yanx27/Pointnet_Pointnet2_pytorch
A pure PyTorch implementation of the PointNet and PointNet++ deep learning architectures for point cloud processing. It includes training a…
324936maintenance
siyuanliii/masa
Official PyTorch implementation of MASA (CVPR 2024 Highlight), a universal instance appearance model that learns to match any objects acros…
331377active
hustvl/VAD
VAD is an end-to-end autonomous driving framework that models the driving scene as a fully vectorized representation of agents and map elem…
601362active
Sense-X/Co-DETR
Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg…
321357stable
ant-research/CoDeF
CoDeF is the official PyTorch implementation of Content Deformation Fields, a video representation combining a canonical content field and …
284846maintenance
mega-sam/mega-sam
MegaSaM is a research codebase implementing a deep visual SLAM system that estimates camera parameters and consistent depth maps from casua…
481355active
lxtGH/OMG-Seg
Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le…
471354active
yinguobing/head-pose-estimation
A Python library for realtime human head pose estimation using ONNX Runtime and OpenCV. It combines face detection (SCRFD), 68-point facial…
231353stable
PKU-VCL-3DV/SLAM3R
SLAM3R is a real-time dense 3D scene reconstruction system that regresses 3D points from monocular RGB video using feed-forward neural netw…
421344active
AI-FanGe/OpenAIglasses_for_Navigation
An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st…
391343active
claritylab/lucida
Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros…
324782maintenance

← prev page 5 / 16 next →