Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
FeiYull/TensorRT-Alpha
A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl…
321460active
dbolya/yolact
YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++…
515241maintenance
zylo117/Yet-Another-EfficientDet-Pytorch
A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim…
235238maintenance
NVIDIA-ISAAC-ROS/isaac_ros_visual_slam
Isaac ROS Visual SLAM is a ROS 2 package providing GPU-accelerated visual simultaneous localization and mapping (VSLAM) using stereo visual…
941448active
amdegroot/ssd.pytorch
A PyTorch implementation of the Single Shot MultiBox Detector (SSD) object detection model from the 2016 paper by Wei Liu et al. It include…
325221maintenance
zsyOAOA/InvSR
InvSR is a Python research library implementing arbitrary-steps image super-resolution via diffusion inversion, leveraging pre-trained diff…
511443active
serratus/quaggaJS
QuaggaJS is a barcode-scanner library written entirely in JavaScript that supports real-time localization and decoding of barcode types suc…
235207maintenance
sdcb/PaddleSharp
A .NET/C# wrapper around Baidu's PaddleInference C API, providing PaddleOCR, PaddleDetection, rotation detection, Chinese segmentation, and…
721441active
valentinfrlch/ha-llmvision
LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live …
901440active
Walter0807/MotionBERT
Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske…
651439active
neuralchen/SimSwap
SimSwap is a PyTorch-based face-swapping framework that performs arbitrary face swaps on images and videos using a single trained model. It…
235188maintenance
Topdu/OpenOCR
OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta…
581437active
jakowenko/double-take
Double Take is a self-hosted Docker application providing a unified UI and API for facial recognition. It abstracts multiple face detection…
401436active
NVlabs/Fast-FoundationStereo
Fast-FoundationStereo is NVIDIA's official PyTorch implementation of a real-time zero-shot stereo matching model family, accepted to CVPR 2…
541432active
Francis-Rings/StableAnimator
StableAnimator is an end-to-end ID-preserving video diffusion framework that animates a reference human image according to a sequence of po…
411430active
jasonmayes/Real-Time-Person-Removal
A browser-based demo that removes people from complex video backgrounds in real time using TensorFlow.js. It learns the static background o…
235154maintenance
zsyOAOA/ResShift
ResShift is an efficient diffusion model for image super-resolution that transfers between low- and high-resolution images by shifting resi…
621427active
zuruoke/watermark-removal
A machine learning tool that removes watermarks from images using deep learning image inpainting, based on Contextual Attention and Gated C…
835139maintenance
tianweiy/CausVid
CausVid is a research codebase implementing a fast autoregressive video diffusion model distilled from a bidirectional diffusion transforme…
361426active
apple/ml-aim
Apple's official repository for AIM (Autoregressive Image Models), providing code and pretrained checkpoints for AIMv1 and AIMv2 large visi…
421424active
ZheC/Realtime_Multi-Person_Pose_Estimation
Reference implementation of the CVPR'17 paper 'Realtime Multi-Person Pose Estimation', a bottom-up approach that detects keypoints for mult…
325123maintenance
facebookresearch/vggsfm
VGGSfM is a deep learning-based Structure from Motion pipeline from Meta AI and Oxford VGG that recovers camera poses and 3D point clouds f…
291421active
SagiPolaczek/NeuralSVG
Official PyTorch implementation of NeuralSVG, an ICCV 2025 paper that generates layered, editable SVG vector graphics from text prompts. It…
471419active
XPandora/PhysGaussian
PhysGaussian is a research library that integrates Material Point Method (MPM) physics simulation with 3D Gaussian Splatting representation…
551414active
CSAILVision/semantic-segmentation-pytorch
A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra…
325078maintenance
khanamiryan/php-qrcode-detector-decoder
A pure PHP library for detecting and decoding QR codes from images, ported from the ZXing library. It works without any PHP extensions beyo…
371412active
mmp/pbrt-v3
pbrt-v3 is the C++ source code for the physically based rendering system described in the third edition of the book 'Physically Based Rende…
325076maintenance
NVlabs/BundleSDF
BundleSDF is a CVPR 2023 research implementation from NVIDIA for near real-time 6-DoF pose tracking of unknown rigid objects from monocular…
651410stable
IDEA-Research/DINO-X-API
DINO-X API is a Python client library and examples for accessing DINO-X, a hosted unified vision model for open-world object detection and …
361410active
nv-tlabs/GEN3C
GEN3C is NVIDIA's research codebase for a generative video model that achieves precise camera control and temporal 3D consistency using a 3…
591409active
dailenson/SDT
Official PyTorch implementation of the CVPR 2023 paper 'Disentangling Writer and Character Styles for Handwriting Generation' (SDT). It gen…
431403active
alibaba/Logics-Parsing
Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul…
541402active
fudan-generative-vision/hallo3
Hallo3 is a research model from Fudan University that animates a single portrait image into a highly dynamic and realistic talking-head vid…
261401active
nachifur/MulimgViewer
MulimgViewer is a Python-based multi-image viewer that displays many images in a single interface for side-by-side comparison, parallel sel…
661400active
imagej/imagej2
ImageJ2 is an open-source Java framework and application for processing and analyzing N-dimensional scientific image data, built on the Img…
751399stable
Zejun-Yang/AniPortrait
AniPortrait is a Python framework from Tencent that generates photorealistic portrait animations from an audio clip and a reference image, …
255021maintenance
zenustech/zeno
ZENO is an open-source, node-based 3D simulation and rendering system written in C++. It lets users build complex physics simulations and v…
671398active
yfeng95/PRNet
PRNet is a Python/TensorFlow implementation of the ECCV 2018 Position Map Regression Network for joint 3D face reconstruction and dense ali…
325013maintenance
om-ai-lab/OmDet
OmDet-Turbo is a PyTorch implementation of a transformer-based open-vocabulary object detection model that detects arbitrary user-defined o…
571393active
zju3dv/street_gaussians
Street Gaussians is a research implementation of the ECCV 2024 paper 'Modeling Dynamic Urban Scenes with Gaussian Splatting', which reconst…
401388active
Junyi42/monst3r
MonST3R is the official PyTorch implementation of an ICLR 2025 paper that estimates per-timestep geometry (pointmaps) from dynamic videos i…
361386active
cszn/BSRGAN
BSRGAN is a PyTorch implementation of a practical degradation model for deep blind image super-resolution, presented at ICCV 2021. It provi…
321386stable
davideberly/GeometricTools
The Geometric Tools Engine (GTE) is a C++14 collection of source code for computing in mathematics, geometry, graphics, image analysis, and…
761382active
zhixuhao/unet
A Keras implementation of the U-Net convolutional network architecture for image segmentation, based on the original biomedical segmentatio…
664941maintenance
autonomousvision/unimatch
UniMatch is a PyTorch research library implementing a unified transformer-based model for optical flow, stereo matching, and depth estimati…
321379stable
yanx27/Pointnet_Pointnet2_pytorch
A pure PyTorch implementation of the PointNet and PointNet++ deep learning architectures for point cloud processing. It includes training a…
324936maintenance
sMythicalBird/ZenlessZoneZero-Auto
A Python-based automation framework for the game Zenless Zone Zero that uses image classification, template matching, and OCR to perform au…
221378active
siyuanliii/masa
Official PyTorch implementation of MASA (CVPR 2024 Highlight), a universal instance appearance model that learns to match any objects acros…
331377active
keyu-tian/SparK
SparK is the official PyTorch implementation of an ICLR 2023 Spotlight paper that applies BERT/MAE-style masked image modeling to convoluti…
221376stable
qubvel/segmentation_models
A Python library providing neural network architectures for image segmentation (Unet, FPN, Linknet, PSPNet) built on Keras and TensorFlow K…
234923maintenance
Meituan-AutoML/MobileVLM
MobileVLM is a family of compact vision language models (1.4B-3B parameters) designed to run efficiently on mobile devices, combining small…
171370active
QwenLM/Qwen3-VL-Embedding
Qwen3-VL-Embedding and Qwen3-VL-Reranker are state-of-the-art multimodal embedding and reranking models built on the Qwen3-VL foundation mo…
561369active
sql-hkr/tiny8
Tiny8 is an educational 8-bit CPU simulator written in Python, featuring an AVR-inspired architecture with 32 registers, a 60+ instruction …
491368active
hustvl/VAD
VAD is an end-to-end autonomous driving framework that models the driving scene as a fully vectorized representation of agents and map elem…
601362active
SBCV/Blender-Addon-Photogrammetry-Importer
A Blender addon that imports photogrammetry and structure-from-motion reconstruction results from tools like COLMAP, Meshroom, OpenMVG, and…
641357active
Sense-X/Co-DETR
Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg…
321357stable
ant-research/CoDeF
CoDeF is the official PyTorch implementation of Content Deformation Fields, a video representation combining a canonical content field and …
284846maintenance
google/neuroglancer
Neuroglancer is a WebGL-based, client-side web application for visualizing large 3-D volumetric datasets, supporting arbitrary cross-sectio…
771355active
mega-sam/mega-sam
MegaSaM is a research codebase implementing a deep visual SLAM system that estimates camera parameters and consistent depth maps from casua…
481355active
lxtGH/OMG-Seg
Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le…
471354active
yinguobing/head-pose-estimation
A Python library for realtime human head pose estimation using ONNX Runtime and OpenCV. It combines face detection (SCRFD), 68-point facial…
231353stable
wyhuai/DDNM
DDNM is a Python research codebase implementing the Denoising Diffusion Null-Space Model for zero-shot image restoration, published as an I…
321349stable
huridocs/pdf-document-layout-analysis
A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl…
821346active
PKU-VCL-3DV/SLAM3R
SLAM3R is a real-time dense 3D scene reconstruction system that regresses 3D points from monocular RGB video using feed-forward neural netw…
421344active
muzishen/IMAGDressing
IMAGDressing-v1 is a diffusion-based framework for customizable virtual dressing that generates human images with fixed garments and contro…
441343active
AI-FanGe/OpenAIglasses_for_Navigation
An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st…
391343active
claritylab/lucida
Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros…
324782maintenance
senguptaumd/Background-Matting
Official research code for 'Background Matting: The World is Your Green Screen' (CVPR 2020), a deep network that extracts per-pixel alpha m…
324769maintenance
IrisRainbowNeko/genshin_auto_fish
A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen…
234758maintenance
wormtql/yas
Yas is a fast screen-scanning tool that uses a custom-trained SVTR OCR model to read Genshin Impact and Honkai: Star Rail artifact stats di…
401336active
nmwsharp/geometry-central
Geometry Central is a modern C++ library of data structures and algorithms for geometry processing, with a particular focus on surface mesh…
771334stable
ByteDance-Seed/SeedVR
SeedVR/SeedVR2 are diffusion-transformer based models for generic real-world and AIGC video and image restoration, with SeedVR2 using adver…
471334active
liwenxi/SWIFT-AI
SWIFT-AI is a deep learning system for extremely fast gigapixel-level visual understanding in scientific applications, such as detecting st…
291334active
mapillary/inplace_abn
A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing…
651333stable
wenqsun/DimensionX
DimensionX is a research framework that generates photorealistic 3D and 4D scenes from a single image using controllable video diffusion mo…
431333active
bytedance/Lance
Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained …
551329active
LLaVA-VL/LLaVA-NeXT
LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision …
644716maintenance
cvzone/cvzone
CVZone is a Python computer vision helper library that wraps OpenCV and MediaPipe to simplify image processing and AI functions like hand t…
321325active
sjtuytc/UnboundedNeRFPytorch
A PyTorch implementation benchmarking state-of-the-art unbounded (large-scale) neural radiance field methods like NeRF++, DVGO, and Block-N…
231324active
galilai-group/lejepa
LeJEPA is a Python framework for scalable, theoretically grounded self-supervised representation learning based on Joint-Embedding Predicti…
451322active
ImprintLab/Medical-SAM-Adapter
Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei…
391322active
neozhaoliang/surround-view-system-introduction
A Python implementation of a vehicle surround-view (bird's-eye view) camera system, covering fisheye camera calibration, projection, image …
661321active
meta-pytorch/segment-anything-fast
A fast, batched offline inference-oriented fork of Meta's Segment Anything (SAM) image segmentation model. It applies optimizations like bf…
451321active
HKUST-Aerial-Robotics/VINS-Fusion
VINS-Fusion is an optimization-based multi-sensor state estimator for accurate self-localization in autonomous applications such as drones,…
324684maintenance
hku-mars/Point-LIO
Point-LIO is a robust high-bandwidth LiDAR-inertial odometry framework that estimates ego-motion and builds maps by fusing LiDAR point clou…
711318active
open-edge-platform/geti
Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru…
981317active
sicara/easy-few-shot-learning
A Python library (easyfsl) with ready-to-use code and tutorial notebooks for few-shot image classification and meta-learning, built on PyTo…
231313stable
ali-vilab/MimicBrush
MimicBrush is the official implementation of a zero-shot image editing method that lets users mask a region in a source image and provide a…
241311active
ndl-lab/ndlocr-lite
NDLOCR-Lite is a lightweight Japanese OCR application developed by the National Diet Library that converts digitized images of books and ma…
771309active
huawei-noah/Efficient-Computing
A collection of efficient deep learning methods from Huawei Noah's Ark Lab, covering model compression, knowledge distillation, pruning, qu…
321307active
DAMO-NLP-SG/VideoLLaMA2
VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro…
251307active
seetaface/SeetaFaceEngine
SeetaFace Engine is an open-source C++ face recognition engine comprising face detection, face alignment, and face identification modules. …
324636maintenance
Vincentqyw/image-matching-webui
A Gradio-based web UI that matches keypoints between two images using many state-of-the-art image matching algorithms (LoFTR, SuperGlue, Li…
911302active
vye16/shape-of-motion
Shape of Motion is a Python research codebase for 4D reconstruction of dynamic scenes from a single monocular video, based on the ICCV 2025…
301302active
OpenTeleVision/TeleVision
Open-TeleVision is an open-source immersive robot teleoperation system that streams stereoscopic visual feedback to VR headsets (Apple Visi…
241301active
STVIR/pysot
PySOT is a Python research platform by SenseTime for single object visual tracking, implementing algorithms such as SiamRPN, SiamRPN++, DaS…
454600maintenance
donydchen/mvsplat
MVSplat is a PyTorch implementation of an ECCV 2024 Oral model that predicts 3D Gaussians from sparse multi-view images in a single feed-fo…
611296active
PyImageSearch/imutils
A Python library of convenience functions that simplify common OpenCV image processing tasks such as translation, rotation, resizing, skele…
324590maintenance
cnr-isti-vclab/vcglib
VCGlib is a templated, header-only C++ library with no external dependencies for manipulating, processing, cleaning, and simplifying triang…
671295stable
fundamentalvision/BEVFormer
BEVFormer is the official PyTorch implementation of an ECCV 2022 paper that learns bird's-eye-view (BEV) representations from multi-camera …
234579maintenance

← prev page 9 / 24 next →