domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Parskatt/RoMa RoMa (romatch) is a Python library for robust dense feature matching between image pairs, estimating pixel-dense warps and reliable certain… | 52 | 1293 | active |
| PantoMatrix/PantoMatrix PantoMatrix is an open-source research project that generates 3D face and body animation from speech audio, including the EMAGE model for h… | 33 | 1293 | active |
| streamlit/demo-self-driving A Streamlit demo app that provides an interactive image browser for the Udacity self-driving-car dataset with realtime YOLO object detectio… | 60 | 1290 | stable |
| RoyalVane/CLAN Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation… | 32 | 1289 | stable |
| hku-mars/livox_camera_calib A C++/ROS tool from HKU MARS for automatic extrinsic calibration between high-resolution LiDAR (e.g., Livox) and cameras in targetless envi… | 32 | 1289 | stable |
| huanngzh/MV-Adapter MV-Adapter is a plug-and-play adapter that turns pre-trained text-to-image diffusion models (e.g., SDXL, SD2.1) into multi-view consistent … | 34 | 1285 | active |
| Phantom-video/HuMo HuMo is a research model and Python codebase from Tsinghua University and ByteDance for human-centric video generation using collaborative … | 46 | 1283 | active |
| BoboTiG/python-mss Python MSS is an ultra-fast, cross-platform screenshot library in pure Python using ctypes, capable of capturing one or all monitors with n… | 81 | 1280 | stable |
| Tianxiaomo/pytorch-YOLOv4 A minimal PyTorch implementation of YOLOv4 (and YOLOv4-tiny) supporting inference and training, with tools to convert Darknet weights to Py… | 32 | 4521 | maintenance |
| zju3dv/MatchAnything MatchAnything is a deep learning model for universal cross-modality image matching, released as research code accompanying a TPAMI 2026 pap… | 64 | 1279 | active |
| AaronJackson/vrn Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses… | 32 | 4517 | maintenance |
| plemeri/transparent-background A Python tool and CLI that removes backgrounds from images and videos using the InSPyReNet deep learning model (ACCV 2022). It supports ima… | 63 | 1278 | active |
| city-super/Scaffold-GS Scaffold-GS is a research implementation of a structured 3D Gaussian splatting method that uses anchor points on a sparse voxel grid to dis… | 27 | 1278 | active |
| nvidia-isaac/nvblox nvblox is a GPU-accelerated C++/Python library for real-time 3D reconstruction using TSDF and ESDF volumetric mapping, designed for robots … | 85 | 1276 | active |
| Fugtemypt123/VIGA VIGA is an analysis-by-synthesis code agent that reconstructs 3D scenes and slide layouts from images by generating and executing Blender P… | 55 | 1275 | active |
| Ma-Lab-Berkeley/CRATE CRATE is the official PyTorch implementation of the Coding RAte reduction TransformEr, a family of 'white-box' transformer architectures de… | 29 | 1275 | active |
| google-research/simclr Google Research's official implementation of SimCLR and SimCLRv2, a framework for contrastive learning of visual representations, with 65 p… | 10 | 4502 | maintenance |
| flutter-ml/google_ml_kit_flutter A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa… | 76 | 1274 | active |
| dcharatan/pixelsplat pixelSplat is a PyTorch implementation of a feed-forward model that reconstructs 3D radiance fields parameterized by 3D Gaussian primitives… | 27 | 1274 | stable |
| nianticlabs/monodepth2 Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra… | 32 | 4497 | maintenance |
| NVlabs/stylegan2-ada-pytorch Official PyTorch implementation of StyleGAN2-ADA, a generative adversarial network with adaptive discriminator augmentation for training wi… | 32 | 4487 | maintenance |
| Visual-Agent/DeepEyes DeepEyes is a research project that trains multimodal vision-language models to 'think with images' using end-to-end reinforcement learning… | 44 | 1271 | active |
| leoxiaobin/deep-high-resolution-net.pytorch Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai… | 32 | 4480 | maintenance |
| AravisProject/aravis Aravis is a C library based on GLib/GObject for video acquisition from Genicam-compliant industrial cameras, implementing GigE Vision and U… | 81 | 1268 | active |
| BachiLi/diffvg diffvg is a differentiable rasterizer for 2D vector graphics that bridges the raster and vector domains via backpropagation. It computes pi… | 42 | 1268 | active |
| amaiya/ktrain ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul… | 25 | 1268 | active |
| ToniRV/NeRF-SLAM NeRF-SLAM is a real-time dense monocular SLAM system that combines neural radiance fields (Instant-NGP) with probabilistic volumetric fusio… | 32 | 1266 | active |
| nv-tlabs/Difix3D Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau… | 32 | 1266 | active |
| bryandlee/animegan2-pytorch A PyTorch implementation of AnimeGANv2, a GAN-based image-to-image style transfer model that converts photos into anime-style images. It pr… | 32 | 4452 | maintenance |
| ethz-asl/rovio ROVIO (Robust Visual Inertial Odometry) is a C++ framework from ETH Zurich that estimates camera and IMU trajectory using an iterated exten… | 61 | 1262 | active |
| rlawjdghek/StableVITON StableVITON is the official PyTorch implementation of a CVPR 2024 paper that performs image-based virtual try-on using a pre-trained latent… | 47 | 1261 | stable |
| Fictionarry/ER-NeRF ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr… | 24 | 1260 | stable |
| nv-tlabs/GET3D GET3D is NVIDIA's PyTorch implementation of a generative model that synthesizes high-quality 3D textured meshes (cars, chairs, animals, bui… | 32 | 4435 | maintenance |
| Yuliang-Liu/MonkeyOCRv2 MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2… | 58 | 1256 | active |
| stardist/stardist StarDist is a Python library for object detection and instance segmentation in 2D and 3D microscopy images using star-convex shapes, built … | 60 | 1255 | stable |
| JonathonLuiten/TrackEval TrackEval is a Python library for evaluating multi-object tracking (MOT) algorithms, implementing metrics such as HOTA, CLEARMOT, IDF1, VAC… | 32 | 1255 | stable |
| ingra14m/Deformable-3D-Gaussians Official PyTorch implementation of the CVPR 2024 paper 'Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction'. … | 18 | 1255 | stable |
| roryclear/clearcam Clearcam is a self-hosted Python NVR that adds AI object detection, tracking, mobile notifications, and semantic search to any RTSP securit… | 86 | 1254 | active |
| ziyc/drivestudio DriveStudio is a Python framework for 3D Gaussian Splatting (3DGS) based reconstruction and simulation of dynamic urban driving scenes. It … | 40 | 1254 | active |
| VainF/pytorch-msssim A PyTorch library providing fast, differentiable SSIM and MS-SSIM image quality metrics using separable Gaussian filtering for speed. It ca… | 23 | 1253 | stable |
| FaceAISDK/FaceAISDK_Android An Android SDK for fully on-device, offline face detection, recognition, liveness detection (anti-spoofing), and 1:1, 1:N, and M:N face sea… | 98 | 1252 | active |
| Linketic/CityGaussian Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia… | 66 | 1251 | active |
| peterbraden/node-opencv Native Node.js bindings for the OpenCV computer vision library, exposing Matrices, image reading/writing, and cascades like face detection … | 23 | 4384 | maintenance |
| withoutbg/withoutbg-python A Python SDK (pip install withoutbg) for removing image backgrounds, offering a free local open-weights ONNX model and an optional paid clo… | 80 | 1246 | active |
| lpiccinelli-eth/UniDepth UniDepth is a Python library and research codebase for universal monocular metric depth estimation from single images, based on CVPR 2024 a… | 35 | 1246 | active |
| unum-cloud/UForm UForm is a compact multimodal AI library providing tiny image-text embedding models (64-768 dimensions, Matryoshka-style) and small generat… | 55 | 1244 | active |
| 3DTopia/OpenLRM OpenLRM is an open-source PyTorch implementation of Large Reconstruction Models (LRM) that reconstruct 3D objects (meshes and rendered vide… | 17 | 1244 | active |
| abewley/sort SORT is a barebones Python implementation of a simple online and realtime multiple object tracking algorithm for 2D video sequences, based … | 32 | 4373 | maintenance |
| HJYao00/Mulberry Mulberry is a research implementation of an o1-like multimodal large language model (MLLM) that performs step-by-step reasoning and reflect… | 49 | 1243 | active |
| cvg/depthsplat DepthSplat is a PyTorch research library implementing a CVPR 2025 model that connects Gaussian splatting with single/multi-view depth estim… | 56 | 1242 | active |
| raspberrypi/picamera2 Picamera2 is a Python library providing an interface to Raspberry Pi cameras via the libcamera stack, replacing the legacy Picamera library… | 94 | 1241 | active |
| alibaba/Tora Tora is Alibaba's official implementation of a trajectory-oriented Diffusion Transformer (DiT) for controllable video generation, integrati… | 64 | 1241 | active |
| XPixelGroup/HYPIR Official PyTorch implementation of HYPIR, a SIGGRAPH 2025 method that harnesses diffusion-yielded score priors for image restoration. It pr… | 39 | 1241 | active |
| jcjohnson/fast-neural-style A Torch (Lua) implementation of feedforward neural style transfer from the ECCV 2016 paper 'Perceptual Losses for Real-Time Style Transfer … | 32 | 4359 | maintenance |
| marian42/mesh_to_sdf A Python library that computes approximate signed distance fields (SDFs) for arbitrary triangle meshes, including non-watertight, self-inte… | 32 | 1240 | stable |
| ZHKKKe/MODNet MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima… | 32 | 4355 | maintenance |
| facebookresearch/deit Official PyTorch repository for DeiT and related vision transformer architectures (CaiT, ResMLP, PatchConvnet, DeiT III), providing trainin… | 10 | 4355 | maintenance |
| apchenstu/TensoRF TensoRF is a PyTorch implementation of the ECCV 2022 paper 'TensoRF: Tensorial Radiance Fields', which models and reconstructs radiance fie… | 44 | 1239 | stable |
| stella-cv/stella_vslam stella_vslam is a community-maintained fork of OpenVSLAM implementing a monocular, stereo, and RGBD visual SLAM system in C++. It supports … | 77 | 1238 | active |
| open-mmlab/playground OpenMMLab Playground is a central hub collecting and showcasing community projects that extend OpenMMLab libraries with Segment Anything Mo… | 30 | 1236 | active |
| bytedance/USO USO is ByteDance's open-source unified style- and subject-driven image generation model based on diffusion (FLUX), combining any subject wi… | 36 | 1235 | active |
| xtreme1-io/xtreme1 Xtreme1 is an open-source, self-hosted data labeling and annotation platform for multimodal training data, supporting images, 3D LiDAR poin… | 62 | 1234 | active |
| facebookresearch/home-robot HomeRobot is an open-source robotics stack from Meta AI for mobile manipulation tasks on low-cost hardware like the Hello Robot Stretch. It… | 23 | 1234 | active |
| VAST-AI-Research/TripoSplat TripoSplat is an inference-only Python library from TripoAI that converts a single 2D image into high-quality 3D Gaussian splats with a var… | 57 | 1232 | active |
| stereolabs/zed-sdk The ZED SDK is a cross-platform spatial perception library for Stereolabs ZED stereo cameras, providing depth sensing, SLAM, 3D reconstruct… | 91 | 1229 | active |
| MoonshotAI/Kimi-VL Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo… | 33 | 1224 | active |
| mrousavy/react-native-fast-tflite A high-performance TensorFlow Lite library for React Native built on Nitro Modules, using the low-level C/C++ TFLite core API with zero-cop… | 84 | 1222 | active |
| maxritter/diy-thermocam DIY-Thermocam is an open-source, self-assembly thermal imaging camera based on the FLIR Lepton sensor and a Teensy 4.1 microcontroller, wit… | 48 | 1222 | active |
| fastgs/FastGS FastGS is a general acceleration framework for 3D Gaussian Splatting that trains scenes in roughly 100 seconds using multi-view consistent … | 49 | 1221 | active |
| MotrixLab/SMPLer-X Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an… | 59 | 1220 | stable |
| bowang-lab/MedRAX MedRAX is a medical reasoning agent framework that integrates chest X-ray analysis tools (segmentation, grounding, report generation, disea… | 42 | 1218 | active |
| fudan-generative-vision/champ Champ is a research framework for controllable and consistent human image animation using 3D parametric guidance (SMPL-based depth, normal,… | 25 | 4261 | maintenance |
| kijai/ComfyUI-segment-anything-2 A set of ComfyUI custom nodes that bring Meta's Segment Anything 2 (SAM2) models into ComfyUI workflows for promptable image and video segm… | 43 | 1214 | active |
| Picsart-AI-Research/Text2Video-Zero Official implementation of Text2Video-Zero, a zero-shot text-to-video generation method that adapts text-to-image diffusion models like Sta… | 30 | 4245 | maintenance |
| ifzhang/FairMOT FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification… | 32 | 4244 | maintenance |
| DachunKai/EvTexture Official PyTorch implementation of EvTexture and EvTexture++, event-driven video super-resolution models that use event-camera signals to e… | 54 | 1207 | active |
| gali8/Tesseract-OCR-iOS An iOS framework wrapping the Tesseract OCR engine (with Leptonica and image libraries) for use in Objective-C or Swift apps on iOS 9.0+. I… | 23 | 4221 | maintenance |
| frotms/PaddleOCR2Pytorch A PyTorch port of PaddleOCR that lets you run PaddleOCR-trained models (detection, recognition, and document structure parsing) without the… | 73 | 1205 | active |
| dexsuite/dex-retargeting A Python library of retargeting optimizers that translate human hand motion (from video or pose datasets) into robot dexterous hand joint c… | 37 | 1205 | active |
| PoseLib/PoseLib PoseLib is a C++ library of minimal solvers for calibrated camera pose estimation, covering absolute and relative pose from point and line … | 68 | 1203 | active |
| dendenxu/fast-gaussian-rasterization A drop-in replacement for diff-gaussian-rasterization that renders 3D Gaussian Splatting scenes using a geometry-shader-based GPU pipeline … | 19 | 1202 | active |
| cleanlab/cleanvision CleanVision is a Python library that automatically detects issues in image datasets, such as blurry, dark, over-exposed, or near-duplicate … | 59 | 1199 | active |
| Calamari-OCR/calamari Calamari is a Python-based OCR engine for line-based automatic text recognition, built on OCRopy and Kraken with a TensorFlow deep-learning… | 74 | 1197 | active |
| toshas/torch-fidelity A PyTorch library providing accurate and efficient implementations of generative model evaluation metrics such as FID, Inception Score, KID… | 70 | 1197 | active |
| deepseek-ai/DeepSeek-VL DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf… | 25 | 4175 | maintenance |
| EvolvingLMMs-Lab/LLaVA-OneVision-2 A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis… | 72 | 1195 | active |
| zai-org/CogAgent CogAgent is an open-source vision-language model (VLM) based GUI agent that understands screen captures and natural language to automate in… | 33 | 1194 | active |
| lessthanoptimal/BoofCV BoofCV is an open-source, real-time computer vision library written entirely in Java, covering image processing, camera calibration, featur… | 86 | 1192 | active |
| Tencent-Hunyuan/HunyuanWorld-Mirror HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds… | 54 | 1191 | active |
| trianglesplatting/triangle-splatting Official implementation of 'Triangle Splatting for Real-Time Radiance Field Rendering' (3DV 2026), which uses 3D triangles as rendering pri… | 43 | 1191 | active |
| shallowdream204/DreamClear DreamClear is a diffusion-transformer based real-world image restoration model for high-fidelity super-resolution, published at NeurIPS 202… | 27 | 1191 | active |
| zai-org/VisualGLM-6B VisualGLM-6B is an open-source multimodal conversational language model supporting images, Chinese, and English, built on ChatGLM-6B with a… | 30 | 4154 | maintenance |
| sair-lab/AirSLAM AirSLAM is an efficient, illumination-robust point-line visual SLAM system supporting stereo visual odometry/VIO, offline map optimization,… | 47 | 1190 | active |
| ali-vilab/UniAnimate UniAnimate is the official code for a research paper on animating a reference human image into a video that follows a driving pose sequence… | 31 | 1189 | active |
| pqpo/SmartCropper An Android library for smart image cropping that automatically detects document borders using OpenCV (with an optional TensorFlow Lite HED … | 66 | 4132 | maintenance |
| SHI-Labs/Neighborhood-Attention-Transformer Official PyTorch implementation of the Neighborhood Attention Transformer (NAT/DiNAT), a family of efficient vision transformers with local… | 32 | 1184 | stable |
| chensjtu/GaussianObject GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,… | 26 | 1183 | active |
| msracver/Deformable-ConvNets Official MXNet implementation of Deformable Convolutional Networks (ICCV 2017) and R-FCN, including deformable convolution and ROI pooling … | 32 | 4121 | maintenance |
| orpatashnik/StyleCLIP Official implementation of StyleCLIP, a method for text-driven manipulation of StyleGAN-generated imagery using CLIP. It provides three app… | 32 | 4121 | maintenance |
| VladimirYugay/Gaussian-SLAM A research implementation of a dense RGBD SLAM system that uses 3D Gaussian Splatting as its scene representation to photorealistically rec… | 27 | 1182 | active |