Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
XPixelGroup/DiffBIR
DiffBIR is a blind image restoration framework that uses generative diffusion priors to restore degraded real-world images. It provides pre…
344119maintenance
mlfoundations/open_flamingo
OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima…
234118maintenance
DAMO-NLP-SG/VideoLLaMA3
VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de…
371179active
ubicomplab/rPPG-Toolbox
rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and…
511178active
NVlabs/Deep_Object_Pose
NVIDIA's Deep Object Pose Estimation (DOPE), a deep learning system for detecting known objects and estimating their 6-DoF pose from RGB ca…
481178active
balancap/SSD-Tensorflow
A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net…
324101maintenance
Soul-AILab/SoulX-LiveAct
SoulX-LiveAct is the official inference code for a real-time human animation framework that generates lifelike, audio/multimodal-controlled…
541176active
tjiiv-cprg/EPro-PnP
EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc…
411175stable
BnanZ0/ok-nte
ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated…
781174active
warmshao/FasterLivePortrait
A real-time portrait animation application based on LivePortrait that animates still photos or videos using a driving video, image, audio, …
381174active
alembic/alembic
Alembic is an open computer graphics interchange framework for storing and sharing baked, animated 3D scene data, consisting of a C++ libra…
841173stable
MIC-DKFZ/batchgenerators
A Python framework for data augmentation of 2D and 3D images, developed by the German Cancer Research Center for medical image classificati…
711173stable
chongzhou96/EdgeSAM
EdgeSAM is the official PyTorch implementation of a distilled, accelerated variant of the Segment Anything Model (SAM) designed for on-devi…
371173active
NVlabs/imaginaire
NVIDIA's PyTorch library containing optimized implementations of image and video synthesis methods, including GAN-based image-to-image tran…
324082maintenance
Kiri-Innovation/3dgs-render-blender-addon
A free, open-source Blender add-on for importing, editing, animating, and rendering 3D Gaussian Splats (3DGS) inside Blender. Created by KI…
901172active
IFL-CAMP/easy_handeye
A ROS package with a GUI that automates hand-eye calibration between a robot and a camera/tracking system, supporting eye-in-hand and eye-o…
471172active
csguoh/MambaIR
MambaIR and MambaIRv2 are PyTorch-based image restoration models built on Mamba state-space models, published at ECCV 2024 and CVPR 2025. T…
541171active
magicleap/SuperGluePretrainedNetwork
SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t…
324072maintenance
juliansteenbakker/mobile_scanner
A Flutter plugin for scanning barcodes and QR codes using the device camera, backed by CameraX/ML Kit on Android, AVFoundation/Apple Vision…
981170active
jbikker/tinybvh
TinyBVH is a single-header, dependency-free C++ library for fast Bounding Volume Hierarchy (BVH) construction and ray traversal on CPU and …
671170active
FoundationVision/GLEE
GLEE is a general object foundation model for images and videos that unifies detection, segmentation, tracking, grounding, and open-world o…
261170active
madmann91/bvh
A header-only C++20 library for constructing and traversing bounding volume hierarchies (BVH), with multiple SAH-based builders, a reinsert…
691169active
jenly1314/MLKit
MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin…
791168active
Object Detection Metrics
A Python toolkit implementing the most popular metrics (AP, mAP, precision-recall curves) used to evaluate object detection algorithms, wit…
501166stable
GaParmar/clean-fid
Clean-FID is a PyTorch library for computing the Frechet Inception Distance (FID) with correct image resizing and quantization steps, fixin…
481166stable
cure-lab/MagicDrive
MagicDrive is the official PyTorch implementation of an ICLR 2024 paper for controllable street view generation using diffusion models with…
351166active
facebookresearch/VideoPose3D
A PyTorch implementation of CVPR 2019 research on 3D human pose estimation in video using temporal convolutions over 2D keypoint trajectori…
104052maintenance
yosinski/deep-visualization-toolbox
A GUI toolbox for visualizing and understanding deep neural networks, showing per-unit activations, backprop/deconv, and regularized-optimi…
324051maintenance
CyberAgentAILab/TANGO
TANGO is a research library from CyberAgent AI Lab that generates co-speech gesture videos by reenactment, using hierarchical audio-motion …
391163active
3D ResNets for Action Recognition
A PyTorch implementation of 3D ResNet and R(2+1)D models for video action recognition, accompanying CVPR 2018 and related papers. It includ…
234038maintenance
Linaom1214/TensorRT-For-YOLO-Series
A Python and C++ toolkit for running YOLO-series object detection models (YOLOv3 through YOLOv12, YOLOX) with NVIDIA TensorRT, including ON…
441162active
MCG-NKU/E2FGVI
E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It…
321161stable
minivision-ai/photo2cartoon
A Python deep-learning project from Minivision that converts real portrait photos into cartoon-style avatars using unpaired image translati…
324029maintenance
DepthAnything/PromptDA
Prompt Depth Anything is a Python library implementing a CVPR 2025 method for high-resolution (up to 4K) accurate metric depth estimation. …
501159active
libuvc/libuvc
libuvc is a cross-platform C library for accessing USB video devices built on top of libusb. It provides fine-grained control over UVC-comp…
321157active
fundamentalvision/Deformable-DETR
Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl…
324015maintenance
sirius-ai/LPRNet_Pytorch
A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus…
321156stable
LazarSoft/jsqrcode
A JavaScript port of the ZXing QR code scanner that decodes QR codes from images or canvas in HTML5-enabled browsers. It supports webcam-ba…
324012maintenance
JunMa11/SegLossOdyssey
A curated collection of loss functions for medical image segmentation, accompanying the 'Loss Odyssey in Medical Image Segmentation' survey…
324007maintenance
SystemErrorWang/White-box-Cartoonization
Official TensorFlow implementation of the CVPR 2020 paper 'Learning to Cartoonize Using White-box Cartoon Representations', which converts …
614001maintenance
OpenGVLab/VisionLLM
VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc…
331153active
quark0/darts
DARTS is the official PyTorch implementation of the ICLR 2019 paper 'DARTS: Differentiable Architecture Search', which performs neural arch…
323997maintenance
DennisLiu1993/Fastest_Image_Pattern_Matching
A C++ library implementing an accelerated Normalized Cross Correlation (NCC)-based template matching and image alignment algorithm, based o…
611151active
mangdangroboticsclub/QuadrupedRobot
Mini Pupper is an open-source ROS-based quadruped robot dog kit built around Raspberry Pi, with software for SLAM, navigation, and OpenCV-b…
671150active
naurril/SUSTechPOINTS
SUSTechPOINTS is a web-based 3D point cloud annotation platform for labeling LiDAR data with 3D bounding boxes, aimed at autonomous driving…
621150active
amazon-science/mm-cot
Official PyTorch implementation of the paper 'Multimodal Chain-of-Thought Reasoning in Language Models', which adds vision features to a tw…
313985maintenance
keith2018/SoftGLRender
A tiny C++ software renderer/rasterizer that emulates the GPU rendering pipeline (vertex/fragment shading, rasterization, depth testing, bl…
741149active
JDAI-CV/fast-reid
FastReID is a PyTorch-based research platform implementing state-of-the-art re-identification algorithms for persons, vehicles, and faces. …
233981maintenance
ttttccxxui/DataInfra-RedactionEverything
A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and pla…
601147active
ShiqiYu/OpenGait
OpenGait is a flexible and extensible Python framework for gait recognition research, providing implementations of state-of-the-art models …
671146active
circuitvalley/USB_C_Industrial_Camera_FPGA_USB3
Open-source USB-C industrial camera project with interchangeable C-mount lens and MIPI sensor, containing PCB designs, Lattice Crosslink NX…
761145active
IDEA-Research/Grounding-DINO-1.5-API
Python examples and API client for Grounding DINO 1.5/1.6, IDEA Research's open-world (open-set) object detection model series hosted on De…
251143active
MIT-SPARK/Hydra
Hydra is a C++ system that incrementally builds hierarchical 3D Scene Graphs from sensor data in real time. It is developed by MIT SPARK as…
681142active
HengyiWang/spann3r
Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.…
261141active
cvg/glue-factory
Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point…
691140active
YuehaiTeam/cocogoat
A browser-based toolbox for Genshin Impact that performs local achievement recognition using PaddleOCR and onnxruntime, plus achievement ma…
761139active
clovaai/deep-text-recognition-benchmark
Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio…
323942maintenance
EyeTrackVR/EyeTrackVR
EyeTrackVR is a free, open-source, DIY software platform that turns affordable cameras and IR LEDs mounted inside a VR headset into an eye …
871138active
chengtan9907/OpenSTL
OpenSTL is a comprehensive benchmark and modular framework for spatio-temporal predictive learning, covering video prediction methods acros…
541137active
facebookresearch/watermark-anything
Official PyTorch implementation and pretrained models for the paper 'Watermark Anything with Localized Messages', which embeds multiple loc…
101137active
kzampog/cilantro
cilantro is a lean, templated C++ library for processing 3D point cloud data, offering kd-trees, normal estimation, resampling, PCA, PLY I/…
451135active
Kiteretsu77/APISR
APISR is a deep-learning based super-resolution tool that restores and enhances low-quality, low-resolution anime images and videos using t…
371135active
SarahWeiii/CoACD
CoACD is a C++ library (with Python bindings and a Unity package) that performs approximate convex decomposition of 3D triangle meshes usin…
981134active
Udayraj123/OMRChecker
OMRChecker is a Python application that reads and evaluates OMR (Optical Mark Recognition) sheets scanned via a scanner or phone camera. It…
671134active
Janspiry/Image-Super-Resolution-via-Iterative-Refinement
An unofficial PyTorch implementation of SR3 (Image Super-Resolution via Iterative Refinement), a diffusion-based model for image super-reso…
323923maintenance
zai-org/SCAIL-2
Official implementation of SCAIL-2, an open-source model for end-to-end controlled character animation that drives character videos from re…
581132active
mgonzs13/yolo_ros
A ROS 2 wrapper for Ultralytics YOLO models (YOLOv8 through YOLO26) providing object detection, tracking, instance segmentation, human pose…
931131active
noahcao/OC_SORT
OC-SORT is a pure motion-model-based multi-object tracker for video, improving on SORT by fixing Kalman filter limitations to handle occlus…
671131stable
yohanshin/WHAM
WHAM is the official PyTorch implementation of the CVPR 2024 paper 'Reconstructing World-grounded Humans with Accurate 3D Motion'. It estim…
261130active
open-mmlab/mmtracking
MMTracking is OpenMMLab's PyTorch-based toolbox for video perception tasks, unifying video object detection, multiple object tracking, sing…
233897maintenance
sml2h3/ddddocr-fastapi
A minimal FastAPI-based REST API service wrapping the DdddOcr OCR engine, exposing endpoints for image text recognition, slide captcha matc…
231126active
brownhci/WebGazer
WebGazer.js is a JavaScript eye tracking library that uses a standard webcam to predict a user's gaze location on a web page in real time. …
653888maintenance
XieZhiFa/IdCardOCR
An Android OCR library for offline recognition of Chinese second-generation ID cards, driver's licenses, and passports. It extracts all fie…
751125active
OpenGVLab/VideoMamba
VideoMamba is a state space model (Mamba-based) architecture for efficient video understanding, released with code and pretrained models fr…
251125active
geopavlakos/hamer
HaMeR (Hand Mesh Recovery) is a transformer-based model that reconstructs 3D hand meshes from single monocular images using the MANO parame…
561124active
xverse-engine/XScene-UEPlugin
An Unreal Engine 5 plugin for real-time visualization, management, editing, and scalable hybrid rendering of 3D Gaussian Splatting models. …
401121active
princeton-vl/RAFT-Stereo
RAFT-Stereo is a PyTorch implementation of a deep learning model for stereo matching that estimates disparity maps from stereo image pairs …
751119stable
caiyuanhao1998/MST
A Python toolbox for spectral compressive imaging reconstruction that implements over 15 algorithms including MST, CST, DAUHST, BiSCI, HDNe…
531118active
yangxue0827/RotationDetection
AlphaRotate is a TensorFlow-based benchmark and toolbox for rotated (oriented) object detection, implementing detectors such as R2CNN, Reti…
231118active
storyicon/comfyui_segment_anything
A ComfyUI custom node that combines GroundingDINO and Segment Anything (SAM) to segment any element in an image using semantic text prompts…
271113active
OpenKinect/libfreenect
libfreenect is a userspace driver and library for the original Microsoft Xbox Kinect sensor, providing access to RGB and depth images, moto…
233826maintenance
Anionex/agent-vision-toolkit
A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, …
791108active
princeton-vl/DPVO
DPVO is a deep learning-based visual odometry and SLAM system that estimates camera trajectories from video or image sequences using patch-…
321108active
THU-MIG/RepViT
Official PyTorch implementation of RepViT, a family of lightweight CNNs designed by integrating efficient ViT architectural designs into Mo…
191108stable
open-mmlab/PowerPaint
PowerPaint is a versatile image inpainting model (ECCV 2024) built on diffusion models that handles text-guided object insertion, object re…
701107active
spytensor/prepare_detection_dataset
A collection of Python scripts that convert object detection datasets between common annotation formats, including CSV, LabelMe JSON, COCO,…
661107active
JEOresearch/EyeTracker
A lightweight open-source Python library for 3D eye tracking that detects and fits the pupil ellipse in eye camera video or images. It is a…
671106active
mlivesu/cinolib
CinoLib is a header-only C++ library for processing polygonal and polyhedral meshes, supporting triangle, quad, and general polygon surface…
671106active
szymanowiczs/splatter-image
Official PyTorch implementation of 'Splatter Image: Ultra-Fast Single-View 3D Reconstruction' (CVPR 2024), which uses an image-to-image net…
261106active
google-research/multinerf
Google Research's official code release for three NeRF papers: Mip-NeRF 360, Ref-NeRF, and RawNeRF, written in JAX. It trains neural radian…
103808maintenance
cameraui/camera.ui
camera.ui is a self-hosted, local-first video surveillance (NVR) platform for security cameras with live viewing, 24/7 recording, and on-de…
1001104active
MIT-SPARK/VGGT-SLAM
VGGT-SLAM is a dense RGB SLAM system that performs real-time feed-forward 3D scene reconstruction, optimizing on the SL(4) manifold using t…
591102active
zclucas/RMT
RMT (RuoMengTu) is a free, open-source macro and desktop automation tool built on AutoHotkey v2. It supports recording and playing keyboard…
881101active
yzslab/gaussian-splatting-lightning
A PyTorch Lightning implementation of 3D Gaussian Splatting with many derived algorithms (Mip-Splatting, LightGaussian, 2DGS, deformable Ga…
581101active
pageauc/speed-camera
A Python3 and OpenCV application that turns a Raspberry Pi, Unix, or Windows computer with a Pi camera, USB webcam, or IP/RTSP camera into …
531101active
superxslam/SuperOdom
SuperOdometry is a lightweight C++/ROS library for LiDAR-only and LiDAR-inertial odometry and mapping, developed by CMU's AirLab. It fuses …
521101active
naturomics/CapsNet-Tensorflow
A TensorFlow implementation of CapsNet (Capsule Networks) based on Geoffrey Hinton's paper 'Dynamic Routing Between Capsules'. It supports …
323786maintenance
WangLibo1995/GeoSeg
GeoSeg is an open-source PyTorch-based semantic segmentation toolbox focused on Vision Transformers for remote sensing imagery, featuring t…
321096active
hkchengrex/Cutie
Cutie is a video object segmentation framework with object-level memory reading, a follow-up to XMem offering better consistency, robustnes…
181095active
LujiaJin/One-Pot_Multi-Frame_Denoising
Official PyTorch implementation of the One-Pot Multi-frame Denoising (OPD) method published at BMVC 2022 and extended in IJCV. It provides …
601094stable

← prev page 11 / 24 next →