domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| BIT-DataLab/Edit-Banana Edit Banana is an open-source Python framework that converts static images and PDFs of diagrams, flowcharts, and charts into fully editable… | 60 | 5469 | active |
| xxlong0/Wonder3D Wonder3D is a cross-domain diffusion model that reconstructs high-fidelity textured 3D meshes from a single image in 2-3 minutes. It genera… | 32 | 5425 | active |
| isl-org/MiDaS MiDaS is a Python library with pretrained models for robust monocular depth estimation from a single image, based on the TPAMI 2022 paper a… | 10 | 5420 | stable |
| facebookresearch/sapiens Sapiens is a family of foundation models from Meta Reality Labs for human-centric vision tasks including 2D pose estimation, body-part segm… | 61 | 5418 | active |
| mayocream/koharu Koharu is a local-first desktop application that automates manga translation using machine learning, combining text/bubble detection, OCR, … | 82 | 5410 | active |
| deepseek-ai/DeepSeek-VL2 DeepSeek-VL2 is a series of Mixture-of-Experts vision-language models (Tiny, Small, and 4.5B activated parameters) with inference code and … | 25 | 5374 | active |
| roboflow/sports A Python library from Roboflow providing reusable computer vision tools for sports analytics, including ball tracking, player tracking and … | 75 | 5320 | active |
| google-ar/arcore-android-sdk Google's ARCore SDK for Android, providing Java and C APIs for building augmented reality experiences with motion tracking, environmental u… | 79 | 5229 | active |
| katanaml/sparrow Sparrow is an open-source framework for structured data extraction from documents (PDFs, images) using ML, LLMs, and Vision LLMs, with sche… | 95 | 5202 | active |
| timesler/facenet-pytorch A PyTorch library providing pretrained face detection (MTCNN) and facial recognition (Inception ResNet V1) models, ported from the TensorFl… | 42 | 5162 | stable |
| NVIDIAGameWorks/kaolin Kaolin is NVIDIA's PyTorch library of GPU-optimized modules for 3D deep learning research, covering meshes, point clouds, and 3D Gaussian s… | 68 | 5161 | active |
| yisol/IDM-VTON Official implementation of IDM-VTON, an ECCV 2024 paper that improves diffusion models for high-fidelity virtual try-on, swapping garments … | 30 | 5156 | active |
| open-mmlab/mmaction2 MMAction2 is OpenMMLab's PyTorch-based toolbox and benchmark for video understanding, covering action recognition, temporal action localiza… | 55 | 5142 | active |
| Breakthrough/PySceneDetect PySceneDetect is a Python and OpenCV-based program and library for detecting scene cuts and transitions in videos, with multiple detection … | 89 | 5123 | stable |
| ai-dawang/PlugNPlay-Modules A curated collection of plug-and-play deep learning modules (convolutions, attention mechanisms, downsampling, and feature fusion blocks) i… | 38 | 5105 | active |
| facebookresearch/AugLy AugLy is a Python data augmentation library from Meta AI supporting audio, image, text, and video with over 100 augmentations. It focuses o… | 67 | 5089 | stable |
| facebookresearch/co-tracker CoTracker is a transformer-based model from Meta AI and Oxford VGG that jointly tracks any point (pixel) across a video, handling occlusion… | 60 | 5080 | active |
| libigl/libigl libigl is a simple C++ geometry processing library offering a wide range of algorithms for triangle and tetrahedral meshes, including defor… | 67 | 5078 | stable |
| opentrack/opentrack opentrack is a head tracking application that captures a user's head movements via webcams, IR trackers, or hardware devices and relays the… | 75 | 5076 | active |
| dmMaze/BallonsTranslator A desktop GUI application that uses deep learning to automatically translate comics and manga, combining text detection, OCR, inpainting, a… | 99 | 5065 | active |
| Deci-AI/super-gradients SuperGradients is an open-source PyTorch-based training library for building, training, and fine-tuning state-of-the-art computer vision mo… | 54 | 5052 | active |
| ArthurBrussee/brush Brush is a 3D reconstruction engine using Gaussian splatting, built in Rust on the Burn ML framework and WebGPU. It trains and renders spla… | 65 | 4992 | active |
| KaiyangZhou/deep-person-reid Torchreid is a PyTorch library for deep-learning person re-identification, supporting both image and video reid with end-to-end training an… | 50 | 4900 | stable |
| tyxsspa/AnyText AnyText is the official implementation of a diffusion-based model for multilingual visual text generation and editing in images, accepted a… | 32 | 4874 | active |
| runhey/OnmyojiAutoScript OnmyojiAutoScript (OAS) is a free, open-source automation script for the mobile game Onmyoji, built on the AzurLaneAutoScript framework. It… | 65 | 4815 | active |
| UX-Decoder/Segment-Everything-Everywhere-All-At-Once SEEM is the official PyTorch implementation of the NeurIPS 2023 paper 'Segment Everything Everywhere All at Once', a model for universal im… | 20 | 4794 | stable |
| zju3dv/EasyMocap EasyMocap is an open-source Python toolbox for markerless human motion capture and novel view synthesis from RGB videos. It fits parametric… | 54 | 4783 | active |
| open-mmlab/mmocr MMOCR is OpenMMLab's PyTorch-based toolbox for text detection, recognition, and key information extraction. It provides a model zoo of OCR … | 23 | 4752 | active |
| OpenDriveLab/UniAD UniAD is a unified end-to-end autonomous driving framework that hierarchically casts perception, prediction, and planning tasks under a pla… | 44 | 4737 | active |
| cvg/LightGlue LightGlue is a deep neural network library that matches sparse local features across image pairs with high accuracy and fast inference. It … | 50 | 4728 | stable |
| esimov/pigo Pigo is a pure Go library for fast face detection, pupil/eye localization, and facial landmark detection based on the Pixel Intensity Compa… | 31 | 4728 | stable |
| MaaXYZ/MaaFramework MaaFramework is an automation black-box testing framework based on image recognition, rewritten from the experience of the MAA (MaaAssistan… | 91 | 4722 | active |
| manycore-research/SpatialLM SpatialLM is a 3D large language model that processes point cloud data (from monocular video, RGBD images, or LiDAR) and generates structur… | 62 | 4719 | active |
| CloudCompare/CloudCompare CloudCompare is a 3D point cloud and triangular mesh processing application, originally built to compare point clouds from laser scanners a… | 67 | 4693 | active |
| f3d-app/f3d F3D is a fast, minimalist open-source 3D viewer desktop application supporting many formats (glTF, USD, STL, STEP, OBJ, FBX, Alembic) with … | 88 | 4650 | active |
| Tencent/TNN TNN is a high-performance, lightweight deep learning inference framework developed by Tencent Youtu Lab, supporting mobile, desktop, and se… | 32 | 4648 | active |
| NVlabs/neuralangelo Official PyTorch implementation of Neuralangelo, a CVPR 2023 method for high-fidelity neural surface reconstruction from multi-view images.… | 29 | 4615 | active |
| sensity-ai/dot dot (Deepfake Offensive Toolkit) is a Python tool that generates real-time, controllable deepfakes from a webcam feed and injects them into… | 23 | 4586 | active |
| joanrod/star-vector StarVector is a foundation model that generates scalable vector graphics (SVG) code from images and text by treating vectorization as a cod… | 49 | 4560 | active |
| hku-mars/FAST-LIVO2 FAST-LIVO2 is a fast, tightly-coupled LiDAR-inertial-visual odometry and mapping system written in C++ on ROS. It provides real-time, accur… | 56 | 4557 | active |
| ceres-solver/ceres-solver Ceres Solver is an open-source C++ library for modeling and solving large-scale non-linear optimization problems, including bounded non-lin… | 76 | 4547 | stable |
| facebookresearch/vjepa2 Official PyTorch codebase and pretrained models for V-JEPA 2, a self-supervised video encoder trained on internet-scale video, plus V-JEPA … | 52 | 4527 | active |
| TencentARC/InstantMesh InstantMesh is a feed-forward framework for generating 3D meshes from a single image using sparse-view large reconstruction models (LRM/Ins… | 25 | 4509 | active |
| royshil/obs-backgroundremoval An OBS Studio plugin that removes and replaces the background in portrait video using ONNX-based machine learning segmentation, acting as a… | 98 | 4492 | active |
| layumi/Person_reID_baseline_pytorch A small, friendly PyTorch baseline implementation for person and vehicle re-identification (ReID). It reproduces strong top-conference resu… | 65 | 4446 | stable |
| rom1504/img2dataset A Python tool that downloads large sets of image URLs and packages them into machine learning datasets, with resizing and caption support. … | 56 | 4443 | active |
| xlite-dev/lite.ai.toolkit A lightweight C++ toolkit providing unified APIs for 100+ pre-trained AI models across inference backends like ONNX Runtime, MNN, TensorRT,… | 74 | 4427 | active |
| huawei-noah/Efficient-AI-Backbones A collection of efficient neural network backbone architectures (GhostNet, TNT, ViG, WaveMLP, TinyNet, etc.) from Huawei Noah's Ark Lab, wi… | 28 | 4418 | active |
| EvolvingLMMs-Lab/lmms-eval lmms-eval is a unified Python framework for evaluating large multimodal models across text, image, video, and audio tasks, with 100+ benchm… | 85 | 4377 | active |
| open-compass/VLMEvalKit VLMEvalKit is an open-source Python toolkit for evaluating large vision-language models (LMMs/LVLMs) across 80+ benchmarks with support for… | 62 | 4359 | active |
| dreamgaussian/dreamgaussian DreamGaussian is the official PyTorch implementation of an ICLR 2024 Oral paper for efficient 3D content creation using generative Gaussian… | 18 | 4352 | active |
| VectorSpaceLab/OmniGen OmniGen is a unified diffusion-based image generation model that produces and edits images from multi-modal prompts without auxiliary modul… | 47 | 4340 | active |
| unum-cloud/USearch USearch is a fast, single-file similarity search and clustering engine for vectors and arbitrary objects, supporting spatial, binary, proba… | 93 | 4278 | active |
| richzhang/PerceptualSimilarity A PyTorch library implementing the LPIPS (Learned Perceptual Image Patch Similarity) metric, which measures perceptual distance between ima… | 23 | 4269 | stable |
| SysCV/sam-hq HQ-SAM (Segment Anything in High Quality) upgrades Meta's Segment Anything Model with a learnable High-Quality Output Token for accurate ze… | 48 | 4255 | active |
| ali-vilab/AnyDoor AnyDoor is the official implementation of a diffusion-based model that teleports target objects into new scenes at user-specified locations… | 28 | 4238 | active |
| cvg/Hierarchical-Localization hloc is a modular Python toolbox for state-of-the-art 6-DoF visual localization, combining image retrieval and feature matching (SuperPoint… | 48 | 4194 | active |
| facebookresearch/vggt-omega VGGT-Omega is a research library from Oxford VGG and Meta AI providing pretrained transformer models for 3D vision tasks such as camera pos… | 58 | 4165 | active |
| torchgeo/torchgeo TorchGeo is a PyTorch domain library, similar to torchvision, providing datasets, samplers, transforms, and pre-trained models specific to … | 95 | 4159 | active |
| justadudewhohacks/face-api.js A JavaScript face detection and face recognition library built on top of tensorflow.js, usable in the browser and Node.js. It provides mode… | 23 | 17945 | maintenance |
| solvespace/solvespace SolveSpace is a free, open-source parametric 2D/3D CAD application with a constraint-based sketcher and solid modeling via extrudes, revolv… | 80 | 4118 | active |
| WebODM/WebODM WebODM is a user-friendly, commercial-grade application for drone image processing that generates georeferenced maps, point clouds, elevati… | 98 | 4116 | active |
| VectorSpaceLab/OmniGen2 OmniGen2 is an open-source unified multimodal generation model supporting text-to-image generation, instruction-guided image editing, and i… | 52 | 4112 | active |
| facebookresearch/jepa Official PyTorch implementation of V-JEPA, a self-supervised method for learning visual representations from video using a joint-embedding … | 29 | 4105 | active |
| cdcseacave/openMVS OpenMVS is an open-source C++ library for Multi-View Stereo 3D reconstruction, taking camera poses and a sparse point-cloud as input and pr… | 76 | 4100 | active |
| ZhengPeng7/BiRefNet BiRefNet is a PyTorch implementation of the CAAI AIR 2024 paper 'Bilateral Reference for High-Resolution Dichotomous Image Segmentation'. I… | 65 | 4098 | active |
| GuyTevet/motion-diffusion-model Official PyTorch implementation of the Human Motion Diffusion Model (MDM) paper, generating 3D human motion sequences from text prompts usi… | 52 | 4092 | active |
| princeton-vl/RAFT Official PyTorch implementation of RAFT (Recurrent All Pairs Field Transforms for Optical Flow), an ECCV 2020 model for estimating dense op… | 49 | 4091 | stable |
| StarsfieldAI/R1-V R1-V is an open-source research codebase for training vision-language models with reinforcement learning (RLVR/GRPO), demonstrating strong … | 21 | 4063 | active |
| Motion-Project/motion Motion is an open-source C++ program that monitors video camera signals and detects changes (motion) in the images. It is commonly used for… | 65 | 4038 | active |
| tensorflow/tensor2tensor Tensor2Tensor (T2T) is a Python library of deep learning models and datasets built on TensorFlow, developed by the Google Brain team to mak… | 10 | 17464 | maintenance |
| cozmo/jsQR jsQR is a pure JavaScript QR code reading library that takes raw image data (RGBA pixel arrays) and locates, extracts, and parses any QR co… | 69 | 4026 | stable |
| introlab/rtabmap RTAB-Map (Real-Time Appearance-Based Mapping) is a C++ library and standalone application implementing graph-based SLAM for RGB-D, stereo, … | 88 | 3965 | stable |
| thuml/Transfer-Learning-Library TLlib is a PyTorch-based open-source library for transfer learning, covering domain adaptation, task adaptation (finetuning), and domain ge… | 23 | 3931 | active |
| hustvl/Vim Vision Mamba (Vim) is a PyTorch implementation of a generic vision backbone built on bidirectional Mamba state space models, published at I… | 29 | 3899 | active |
| hustvl/4DGaussians An official PyTorch implementation of 4D Gaussian Splatting (4D-GS) for real-time rendering of dynamic scenes, published at CVPR 2024. It c… | 27 | 3895 | active |
| tinyobjloader/tinyobjloader A tiny, dependency-free Wavefront .obj/.mtl parser available as a single-header C++11 library and a pure C11 implementation with polygon te… | 62 | 3866 | active |
| NVlabs/VILA VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d… | 57 | 3857 | active |
| mseitzer/pytorch-fid A PyTorch port of the official TensorFlow implementation of the Fréchet Inception Distance (FID), a metric for measuring similarity between… | 23 | 3851 | stable |
| open-mmlab/mmpretrain MMPretrain is OpenMMLab's PyTorch-based toolbox and benchmark for image classification model pre-training, covering supervised, self-superv… | 23 | 3850 | active |
| Avatarify Avatarify is an open-source application that drives photorealistic avatars in real time for video-conferencing apps like Zoom and Skype, ba… | 23 | 16515 | maintenance |
| google-research/scenic Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr… | 76 | 3821 | active |
| CHNZYX/Auto_Simulated_Universe A Python-based automation tool for the Honkai: Star Rail 'Simulated Universe' game mode, using screen recognition to play the roguelike mod… | 87 | 3811 | active |
| lightly-ai/lightly LightlySSL is a Python library built on PyTorch for self-supervised learning on images, offering modular implementations of methods like Si… | 93 | 3797 | active |
| fudan-generative-vision/hallo2 Hallo2 is a Python research library from Fudan University that animates a single portrait image using audio input, producing long-duration … | 26 | 3734 | active |
| abhiTronix/vidgear VidGear is a high-performance, cross-platform Python framework for video processing built around multi-threaded and asynchronous pipelines.… | 76 | 3721 | active |
| roboflow/trackers A Python library of clean-room, Apache 2.0 implementations of multi-object tracking algorithms including SORT, ByteTrack, OC-SORT, BoT-SORT… | 85 | 3717 | active |
| MaaEnd/MaaEnd MaaEnd is a vision-AI-powered automation assistant for the game 'Arknights: Endfield', built on MaaFramework. It captures the screen, recog… | 95 | 3712 | active |
| IDEA-Research/Grounded-SAM-2 Grounded SAM 2 is a foundation-model pipeline that combines open-set detectors (Grounding DINO, Grounding DINO 1.5/1.6, Florence-2, DINO-X)… | 37 | 3708 | active |
| xinyu1205/recognize-anything Recognize Anything is a collection of open-source image recognition foundation models, including RAM, RAM++, and Tag2Text, that perform ima… | 33 | 3708 | active |
| Belval/TextRecognitionDataGenerator A Python library and CLI tool (trdg) that generates synthetic text images for training OCR and text recognition models. It supports multipl… | 23 | 3691 | stable |
| ant-research/MagicQuill MagicQuill is an intelligent interactive image editing system from a CVPR 2025 paper, combining a brush-based UI with AI-powered suggestion… | 46 | 3688 | active |
| microsoft/Bringing-Old-Photos-Back-to-Life The official PyTorch implementation of 'Bringing Old Photos Back to Life' (CVPR 2020 Oral), a deep learning model that restores old photos … | 23 | 15704 | maintenance |
| DLR-RM/BlenderProc BlenderProc is a procedural Python pipeline built on Blender for generating photorealistic synthetic training images with ground-truth anno… | 62 | 3684 | active |
| mmp/pbrt-v4 pbrt-v4 is the C++ physically based ray tracing system accompanying the fourth edition of the book 'Physically Based Rendering: From Theory… | 71 | 3683 | stable |
| stack-of-tasks/pinocchio Pinocchio is a fast C++ library (with Python bindings) implementing state-of-the-art rigid body dynamics algorithms for poly-articulated sy… | 93 | 3682 | active |
| facebookresearch/map-anything MapAnything is an open-source research framework from Meta and CMU for universal feed-forward metric 3D reconstruction using an end-to-end … | 77 | 3682 | active |
| ferdous-alam/GenCAD GenCAD is a research codebase for image-conditioned CAD model generation using transformer-based contrastive representations (CCIP) and dif… | 36 | 3669 | active |
| mikedh/trimesh Trimesh is a pure Python library for loading, manipulating, and analyzing triangular meshes with an emphasis on watertight surfaces. It pro… | 98 | 3660 | stable |
| borglab/gtsam GTSAM is a C++ library implementing smoothing and mapping (SAM) for robotics and vision using factor graphs and Bayes networks as its core … | 92 | 3656 | active |