domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| FeiYull/TensorRT-Alpha A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl… | 32 | 1460 | active |
| dbolya/yolact YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++… | 51 | 5241 | maintenance |
| zylo117/Yet-Another-EfficientDet-Pytorch A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim… | 23 | 5238 | maintenance |
| NVIDIA-ISAAC-ROS/isaac_ros_visual_slam Isaac ROS Visual SLAM is a ROS 2 package providing GPU-accelerated visual simultaneous localization and mapping (VSLAM) using stereo visual… | 94 | 1448 | active |
| amdegroot/ssd.pytorch A PyTorch implementation of the Single Shot MultiBox Detector (SSD) object detection model from the 2016 paper by Wei Liu et al. It include… | 32 | 5221 | maintenance |
| zsyOAOA/InvSR InvSR is a Python research library implementing arbitrary-steps image super-resolution via diffusion inversion, leveraging pre-trained diff… | 51 | 1443 | active |
| serratus/quaggaJS QuaggaJS is a barcode-scanner library written entirely in JavaScript that supports real-time localization and decoding of barcode types suc… | 23 | 5207 | maintenance |
| sdcb/PaddleSharp A .NET/C# wrapper around Baidu's PaddleInference C API, providing PaddleOCR, PaddleDetection, rotation detection, Chinese segmentation, and… | 72 | 1441 | active |
| valentinfrlch/ha-llmvision LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live … | 90 | 1440 | active |
| Walter0807/MotionBERT Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske… | 65 | 1439 | active |
| neuralchen/SimSwap SimSwap is a PyTorch-based face-swapping framework that performs arbitrary face swaps on images and videos using a single trained model. It… | 23 | 5188 | maintenance |
| Topdu/OpenOCR OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta… | 58 | 1437 | active |
| jakowenko/double-take Double Take is a self-hosted Docker application providing a unified UI and API for facial recognition. It abstracts multiple face detection… | 40 | 1436 | active |
| NVlabs/Fast-FoundationStereo Fast-FoundationStereo is NVIDIA's official PyTorch implementation of a real-time zero-shot stereo matching model family, accepted to CVPR 2… | 54 | 1432 | active |
| Francis-Rings/StableAnimator StableAnimator is an end-to-end ID-preserving video diffusion framework that animates a reference human image according to a sequence of po… | 41 | 1430 | active |
| jasonmayes/Real-Time-Person-Removal A browser-based demo that removes people from complex video backgrounds in real time using TensorFlow.js. It learns the static background o… | 23 | 5154 | maintenance |
| zsyOAOA/ResShift ResShift is an efficient diffusion model for image super-resolution that transfers between low- and high-resolution images by shifting resi… | 62 | 1427 | active |
| zuruoke/watermark-removal A machine learning tool that removes watermarks from images using deep learning image inpainting, based on Contextual Attention and Gated C… | 83 | 5139 | maintenance |
| tianweiy/CausVid CausVid is a research codebase implementing a fast autoregressive video diffusion model distilled from a bidirectional diffusion transforme… | 36 | 1426 | active |
| apple/ml-aim Apple's official repository for AIM (Autoregressive Image Models), providing code and pretrained checkpoints for AIMv1 and AIMv2 large visi… | 42 | 1424 | active |
| ZheC/Realtime_Multi-Person_Pose_Estimation Reference implementation of the CVPR'17 paper 'Realtime Multi-Person Pose Estimation', a bottom-up approach that detects keypoints for mult… | 32 | 5123 | maintenance |
| facebookresearch/vggsfm VGGSfM is a deep learning-based Structure from Motion pipeline from Meta AI and Oxford VGG that recovers camera poses and 3D point clouds f… | 29 | 1421 | active |
| SagiPolaczek/NeuralSVG Official PyTorch implementation of NeuralSVG, an ICCV 2025 paper that generates layered, editable SVG vector graphics from text prompts. It… | 47 | 1419 | active |
| XPandora/PhysGaussian PhysGaussian is a research library that integrates Material Point Method (MPM) physics simulation with 3D Gaussian Splatting representation… | 55 | 1414 | active |
| CSAILVision/semantic-segmentation-pytorch A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra… | 32 | 5078 | maintenance |
| khanamiryan/php-qrcode-detector-decoder A pure PHP library for detecting and decoding QR codes from images, ported from the ZXing library. It works without any PHP extensions beyo… | 37 | 1412 | active |
| mmp/pbrt-v3 pbrt-v3 is the C++ source code for the physically based rendering system described in the third edition of the book 'Physically Based Rende… | 32 | 5076 | maintenance |
| NVlabs/BundleSDF BundleSDF is a CVPR 2023 research implementation from NVIDIA for near real-time 6-DoF pose tracking of unknown rigid objects from monocular… | 65 | 1410 | stable |
| IDEA-Research/DINO-X-API DINO-X API is a Python client library and examples for accessing DINO-X, a hosted unified vision model for open-world object detection and … | 36 | 1410 | active |
| nv-tlabs/GEN3C GEN3C is NVIDIA's research codebase for a generative video model that achieves precise camera control and temporal 3D consistency using a 3… | 59 | 1409 | active |
| dailenson/SDT Official PyTorch implementation of the CVPR 2023 paper 'Disentangling Writer and Character Styles for Handwriting Generation' (SDT). It gen… | 43 | 1403 | active |
| alibaba/Logics-Parsing Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul… | 54 | 1402 | active |
| fudan-generative-vision/hallo3 Hallo3 is a research model from Fudan University that animates a single portrait image into a highly dynamic and realistic talking-head vid… | 26 | 1401 | active |
| nachifur/MulimgViewer MulimgViewer is a Python-based multi-image viewer that displays many images in a single interface for side-by-side comparison, parallel sel… | 66 | 1400 | active |
| imagej/imagej2 ImageJ2 is an open-source Java framework and application for processing and analyzing N-dimensional scientific image data, built on the Img… | 75 | 1399 | stable |
| Zejun-Yang/AniPortrait AniPortrait is a Python framework from Tencent that generates photorealistic portrait animations from an audio clip and a reference image, … | 25 | 5021 | maintenance |
| zenustech/zeno ZENO is an open-source, node-based 3D simulation and rendering system written in C++. It lets users build complex physics simulations and v… | 67 | 1398 | active |
| yfeng95/PRNet PRNet is a Python/TensorFlow implementation of the ECCV 2018 Position Map Regression Network for joint 3D face reconstruction and dense ali… | 32 | 5013 | maintenance |
| om-ai-lab/OmDet OmDet-Turbo is a PyTorch implementation of a transformer-based open-vocabulary object detection model that detects arbitrary user-defined o… | 57 | 1393 | active |
| zju3dv/street_gaussians Street Gaussians is a research implementation of the ECCV 2024 paper 'Modeling Dynamic Urban Scenes with Gaussian Splatting', which reconst… | 40 | 1388 | active |
| Junyi42/monst3r MonST3R is the official PyTorch implementation of an ICLR 2025 paper that estimates per-timestep geometry (pointmaps) from dynamic videos i… | 36 | 1386 | active |
| cszn/BSRGAN BSRGAN is a PyTorch implementation of a practical degradation model for deep blind image super-resolution, presented at ICCV 2021. It provi… | 32 | 1386 | stable |
| davideberly/GeometricTools The Geometric Tools Engine (GTE) is a C++14 collection of source code for computing in mathematics, geometry, graphics, image analysis, and… | 76 | 1382 | active |
| zhixuhao/unet A Keras implementation of the U-Net convolutional network architecture for image segmentation, based on the original biomedical segmentatio… | 66 | 4941 | maintenance |
| autonomousvision/unimatch UniMatch is a PyTorch research library implementing a unified transformer-based model for optical flow, stereo matching, and depth estimati… | 32 | 1379 | stable |
| yanx27/Pointnet_Pointnet2_pytorch A pure PyTorch implementation of the PointNet and PointNet++ deep learning architectures for point cloud processing. It includes training a… | 32 | 4936 | maintenance |
| sMythicalBird/ZenlessZoneZero-Auto A Python-based automation framework for the game Zenless Zone Zero that uses image classification, template matching, and OCR to perform au… | 22 | 1378 | active |
| siyuanliii/masa Official PyTorch implementation of MASA (CVPR 2024 Highlight), a universal instance appearance model that learns to match any objects acros… | 33 | 1377 | active |
| keyu-tian/SparK SparK is the official PyTorch implementation of an ICLR 2023 Spotlight paper that applies BERT/MAE-style masked image modeling to convoluti… | 22 | 1376 | stable |
| qubvel/segmentation_models A Python library providing neural network architectures for image segmentation (Unet, FPN, Linknet, PSPNet) built on Keras and TensorFlow K… | 23 | 4923 | maintenance |
| Meituan-AutoML/MobileVLM MobileVLM is a family of compact vision language models (1.4B-3B parameters) designed to run efficiently on mobile devices, combining small… | 17 | 1370 | active |
| QwenLM/Qwen3-VL-Embedding Qwen3-VL-Embedding and Qwen3-VL-Reranker are state-of-the-art multimodal embedding and reranking models built on the Qwen3-VL foundation mo… | 56 | 1369 | active |
| sql-hkr/tiny8 Tiny8 is an educational 8-bit CPU simulator written in Python, featuring an AVR-inspired architecture with 32 registers, a 60+ instruction … | 49 | 1368 | active |
| hustvl/VAD VAD is an end-to-end autonomous driving framework that models the driving scene as a fully vectorized representation of agents and map elem… | 60 | 1362 | active |
| SBCV/Blender-Addon-Photogrammetry-Importer A Blender addon that imports photogrammetry and structure-from-motion reconstruction results from tools like COLMAP, Meshroom, OpenMVG, and… | 64 | 1357 | active |
| Sense-X/Co-DETR Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg… | 32 | 1357 | stable |
| ant-research/CoDeF CoDeF is the official PyTorch implementation of Content Deformation Fields, a video representation combining a canonical content field and … | 28 | 4846 | maintenance |
| google/neuroglancer Neuroglancer is a WebGL-based, client-side web application for visualizing large 3-D volumetric datasets, supporting arbitrary cross-sectio… | 77 | 1355 | active |
| mega-sam/mega-sam MegaSaM is a research codebase implementing a deep visual SLAM system that estimates camera parameters and consistent depth maps from casua… | 48 | 1355 | active |
| lxtGH/OMG-Seg Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le… | 47 | 1354 | active |
| yinguobing/head-pose-estimation A Python library for realtime human head pose estimation using ONNX Runtime and OpenCV. It combines face detection (SCRFD), 68-point facial… | 23 | 1353 | stable |
| wyhuai/DDNM DDNM is a Python research codebase implementing the Denoising Diffusion Null-Space Model for zero-shot image restoration, published as an I… | 32 | 1349 | stable |
| huridocs/pdf-document-layout-analysis A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl… | 82 | 1346 | active |
| PKU-VCL-3DV/SLAM3R SLAM3R is a real-time dense 3D scene reconstruction system that regresses 3D points from monocular RGB video using feed-forward neural netw… | 42 | 1344 | active |
| muzishen/IMAGDressing IMAGDressing-v1 is a diffusion-based framework for customizable virtual dressing that generates human images with fixed garments and contro… | 44 | 1343 | active |
| AI-FanGe/OpenAIglasses_for_Navigation An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st… | 39 | 1343 | active |
| claritylab/lucida Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros… | 32 | 4782 | maintenance |
| senguptaumd/Background-Matting Official research code for 'Background Matting: The World is Your Green Screen' (CVPR 2020), a deep network that extracts per-pixel alpha m… | 32 | 4769 | maintenance |
| IrisRainbowNeko/genshin_auto_fish A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen… | 23 | 4758 | maintenance |
| wormtql/yas Yas is a fast screen-scanning tool that uses a custom-trained SVTR OCR model to read Genshin Impact and Honkai: Star Rail artifact stats di… | 40 | 1336 | active |
| nmwsharp/geometry-central Geometry Central is a modern C++ library of data structures and algorithms for geometry processing, with a particular focus on surface mesh… | 77 | 1334 | stable |
| ByteDance-Seed/SeedVR SeedVR/SeedVR2 are diffusion-transformer based models for generic real-world and AIGC video and image restoration, with SeedVR2 using adver… | 47 | 1334 | active |
| liwenxi/SWIFT-AI SWIFT-AI is a deep learning system for extremely fast gigapixel-level visual understanding in scientific applications, such as detecting st… | 29 | 1334 | active |
| mapillary/inplace_abn A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing… | 65 | 1333 | stable |
| wenqsun/DimensionX DimensionX is a research framework that generates photorealistic 3D and 4D scenes from a single image using controllable video diffusion mo… | 43 | 1333 | active |
| bytedance/Lance Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained … | 55 | 1329 | active |
| LLaVA-VL/LLaVA-NeXT LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision … | 64 | 4716 | maintenance |
| cvzone/cvzone CVZone is a Python computer vision helper library that wraps OpenCV and MediaPipe to simplify image processing and AI functions like hand t… | 32 | 1325 | active |
| sjtuytc/UnboundedNeRFPytorch A PyTorch implementation benchmarking state-of-the-art unbounded (large-scale) neural radiance field methods like NeRF++, DVGO, and Block-N… | 23 | 1324 | active |
| galilai-group/lejepa LeJEPA is a Python framework for scalable, theoretically grounded self-supervised representation learning based on Joint-Embedding Predicti… | 45 | 1322 | active |
| ImprintLab/Medical-SAM-Adapter Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei… | 39 | 1322 | active |
| neozhaoliang/surround-view-system-introduction A Python implementation of a vehicle surround-view (bird's-eye view) camera system, covering fisheye camera calibration, projection, image … | 66 | 1321 | active |
| meta-pytorch/segment-anything-fast A fast, batched offline inference-oriented fork of Meta's Segment Anything (SAM) image segmentation model. It applies optimizations like bf… | 45 | 1321 | active |
| HKUST-Aerial-Robotics/VINS-Fusion VINS-Fusion is an optimization-based multi-sensor state estimator for accurate self-localization in autonomous applications such as drones,… | 32 | 4684 | maintenance |
| hku-mars/Point-LIO Point-LIO is a robust high-bandwidth LiDAR-inertial odometry framework that estimates ego-motion and builds maps by fusing LiDAR point clou… | 71 | 1318 | active |
| open-edge-platform/geti Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru… | 98 | 1317 | active |
| sicara/easy-few-shot-learning A Python library (easyfsl) with ready-to-use code and tutorial notebooks for few-shot image classification and meta-learning, built on PyTo… | 23 | 1313 | stable |
| ali-vilab/MimicBrush MimicBrush is the official implementation of a zero-shot image editing method that lets users mask a region in a source image and provide a… | 24 | 1311 | active |
| ndl-lab/ndlocr-lite NDLOCR-Lite is a lightweight Japanese OCR application developed by the National Diet Library that converts digitized images of books and ma… | 77 | 1309 | active |
| huawei-noah/Efficient-Computing A collection of efficient deep learning methods from Huawei Noah's Ark Lab, covering model compression, knowledge distillation, pruning, qu… | 32 | 1307 | active |
| DAMO-NLP-SG/VideoLLaMA2 VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro… | 25 | 1307 | active |
| seetaface/SeetaFaceEngine SeetaFace Engine is an open-source C++ face recognition engine comprising face detection, face alignment, and face identification modules. … | 32 | 4636 | maintenance |
| Vincentqyw/image-matching-webui A Gradio-based web UI that matches keypoints between two images using many state-of-the-art image matching algorithms (LoFTR, SuperGlue, Li… | 91 | 1302 | active |
| vye16/shape-of-motion Shape of Motion is a Python research codebase for 4D reconstruction of dynamic scenes from a single monocular video, based on the ICCV 2025… | 30 | 1302 | active |
| OpenTeleVision/TeleVision Open-TeleVision is an open-source immersive robot teleoperation system that streams stereoscopic visual feedback to VR headsets (Apple Visi… | 24 | 1301 | active |
| STVIR/pysot PySOT is a Python research platform by SenseTime for single object visual tracking, implementing algorithms such as SiamRPN, SiamRPN++, DaS… | 45 | 4600 | maintenance |
| donydchen/mvsplat MVSplat is a PyTorch implementation of an ECCV 2024 Oral model that predicts 3D Gaussians from sparse multi-view images in a single feed-fo… | 61 | 1296 | active |
| PyImageSearch/imutils A Python library of convenience functions that simplify common OpenCV image processing tasks such as translation, rotation, resizing, skele… | 32 | 4590 | maintenance |
| cnr-isti-vclab/vcglib VCGlib is a templated, header-only C++ library with no external dependencies for manipulating, processing, cleaning, and simplifying triang… | 67 | 1295 | stable |
| fundamentalvision/BEVFormer BEVFormer is the official PyTorch implementation of an ECCV 2022 paper that learns bird's-eye-view (BEV) representations from multi-camera … | 23 | 4579 | maintenance |