domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| open-mmlab/mmsegmentation MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat… | 23 | 9930 | stable |
| playcanvas/supersplat SuperSplat is a free, open-source browser-based editor for inspecting, editing, optimizing, and publishing 3D Gaussian Splats. It is built … | 90 | 9904 | active |
| StarTrail-org/PixelRAG PixelRAG is a Python library and hosted service for visual retrieval-augmented generation: it renders web pages and documents into screensh… | 76 | 9742 | active |
| mrousavy/react-native-vision-camera A high-performance camera library for React Native offering photo/video capture, QR/barcode scanning, and JS worklet-based frame processors… | 99 | 9580 | active |
| WongKinYiu/yolov9 Official PyTorch implementation of the YOLOv9 object detection paper, featuring Programmable Gradient Information for improved accuracy. It… | 16 | 9551 | active |
| PeterL1n/RobustVideoMatting Robust Video Matting (RVM) is a deep learning model and library for real-time human video matting, using a recurrent neural network with te… | 23 | 9500 | stable |
| PaddlePaddle/PaddleSeg PaddleSeg is an end-to-end image segmentation toolkit built on PaddlePaddle, offering a model zoo with dozens of pre-trained models for sem… | 52 | 9382 | active |
| lipku/LiveTalking LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit… | 89 | 9238 | active |
| studio-dots-ai/dots.ocr dots.ocr is a 1.7B-parameter vision-language model for multilingual document layout parsing, converting documents into structured output wi… | 51 | 9090 | active |
| roboflow/rf-detr RF-DETR is a real-time transformer-based model architecture from Roboflow for object detection, instance segmentation, and keypoint detecti… | 87 | 9063 | active |
| bytedance/Dolphin Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-… | 52 | 9049 | active |
| RealSense SDK RealSense SDK 2.0 (librealsense) is a cross-platform C++ library for Intel/RealSense depth cameras, providing depth and color streaming plu… | 93 | 8974 | active |
| dusty-nv/jetson-inference A C++/Python DNN inference library and tutorial guide ('Hello AI World') for deploying deep learning vision models on NVIDIA Jetson devices… | 44 | 8969 | stable |
| apple/ml-sharp SHARP is a Python tool from Apple that synthesizes a photorealistic 3D Gaussian splat representation from a single photograph in under a se… | 42 | 8843 | active |
| PantsuDango/Dango-Translator Dango-Translator (团子翻译器) is a Windows desktop application that performs real-time OCR-based translation of on-screen text ('raw' untranslat… | 91 | 8751 | active |
| FoundationVision/VAR Official PyTorch implementation of Visual Autoregressive Modeling (VAR), a NeurIPS 2024 Best Paper-winning method for scalable image genera… | 48 | 8729 | active |
| DepthAnything/Depth-Anything-V2 Depth Anything V2 is a foundation model for monocular depth estimation, trained on 595K synthetic labeled images and 62M+ real unlabeled im… | 56 | 8709 | stable |
| fudan-generative-vision/hallo Hallo is a Python research library implementing hierarchical audio-driven visual synthesis for animating portrait images into talking-head … | 14 | 8664 | active |
| jomjol/AI-on-the-edge-device A firmware application for ESP32-CAM boards that uses TensorFlow Lite CNNs on-device to digitize analog utility meters (water, gas, electri… | 72 | 8624 | active |
| CASIA-LMC-Lab/FastSAM FastSAM is a CNN-based Segment Anything Model trained on only 2% of the SA-1B dataset, achieving comparable segmentation performance to SAM… | 19 | 8401 | active |
| XPixelGroup/BasicSR BasicSR is an open-source PyTorch toolbox for image and video restoration tasks such as super-resolution, denoising, deblurring, and JPEG a… | 23 | 8367 | stable |
| bytedeco/javacv JavaCV is a Java library that wraps OpenCV, FFmpeg, and other computer vision and multimedia libraries via JavaCPP Presets, with utility cl… | 86 | 8335 | active |
| mikel-brostrom/boxmot BoxMOT is a pluggable Python and C++ library providing state-of-the-art multi-object tracking (MOT) algorithms such as ByteTrack, BoT-SORT,… | 95 | 8281 | active |
| exadel-inc/CompreFace Exadel CompreFace is a free, open-source face recognition system that provides REST APIs for face recognition, verification, detection, lan… | 23 | 8273 | stable |
| Ucas-HaoranWei/GOT-OCR2.0 Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, … | 25 | 8216 | active |
| GetStream/Vision-Agents An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O… | 84 | 8100 | active |
| bingoogolapple/BGAQRCode-Android An Android library for scanning and generating QR codes and barcodes, offering both ZXing and ZBar engines behind customizable scan views. … | 64 | 8008 | stable |
| open-mmlab/mmpose MMPose is an open-source pose estimation toolbox and benchmark built on PyTorch as part of the OpenMMLab ecosystem. It provides implementat… | 39 | 7855 | active |
| wang-xinyu/tensorrtx A C++ collection of popular deep learning networks (YOLO variants, ResNet, MobileNet, Swin Transformer, OCR models, and more) implemented f… | 76 | 7827 | active |
| TencentARC/GFPGAN GFPGAN is a Python library built on PyTorch that restores and enhances real-world degraded face photos using GAN-based priors. It provides … | 23 | 37657 | maintenance |
| TadasBaltrusaitis/OpenFace OpenFace is a facial behavior analysis toolkit that performs facial landmark detection, head pose estimation, facial action unit recognitio… | 23 | 7740 | active |
| boltgolt/howdy Howdy provides Windows Hello-style facial authentication for Linux using IR emitters and a camera with facial recognition. It integrates wi… | 38 | 7738 | active |
| geekyutao/Inpaint-Anything Inpaint Anything combines Segment Anything (SAM) with inpainting models like LaMa and Stable Diffusion to remove, fill, or replace objects … | 65 | 7703 | active |
| RapidAI/RapidOCR RapidOCR is an open-source, multi-language OCR toolkit that performs text detection and recognition using models converted to run on ONNX R… | 97 | 7599 | active |
| 1adrianb/face-alignment A Python library built on PyTorch that detects 2D and 3D facial landmarks in images using the FAN deep learning face alignment network. It … | 70 | 7538 | active |
| hybridgroup/gocv GoCV is a Go language binding for the OpenCV 4 computer vision library, supporting Linux, macOS, Windows, and Docker. It includes support f… | 76 | 7491 | active |
| google/draco Draco is an open-source C++ library from Google for compressing and decompressing 3D geometric meshes and point clouds. It improves the sto… | 67 | 7458 | stable |
| apple/ml-fastvlm Official implementation of FastVLM, a vision language model with an efficient hybrid vision encoder (FastViTHD) that reduces token count an… | 28 | 7411 | active |
| facebookresearch/SlowFast PySlowFast is a PyTorch-based open-source video understanding codebase from Facebook AI Research (FAIR). It provides implementations of sta… | 65 | 7410 | active |
| zai-org/GLM-OCR GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understa… | 65 | 7366 | active |
| EutropicAI/Final2x Final2x is a cross-platform desktop application for image super-resolution (upscaling) built with Electron, Vue3, and a PyTorch-based Pytho… | 74 | 7323 | active |
| facebookresearch/sam-3d-objects SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code… | 55 | 7322 | active |
| naver/dust3r DUSt3R is the official PyTorch implementation of a CVPR 2024 model that performs dense, unconstrained stereo and multi-view 3D reconstructi… | 45 | 7288 | active |
| liuliu/ccv ccv is a modern, minimalist computer vision library written in C/C++ with an application-driven set of state-of-the-art algorithms includin… | 77 | 7243 | active |
| princeton-vl/infinigen Infinigen is a procedural generator of infinite photorealistic 3D worlds and scenes, built on Blender by the Princeton Vision & Learning La… | 94 | 7222 | active |
| BVLC/caffe Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp… | 23 | 34556 | maintenance |
| PeterL1n/BackgroundMattingV2 Official PyTorch implementation of the CVPR 2021 paper 'Real-Time High-Resolution Background Matting'. It produces state-of-the-art alpha m… | 23 | 7189 | stable |
| CMU-Perceptual-Computing-Lab/openpose OpenPose is a real-time multi-person keypoint detection library that jointly estimates human body, face, hand, and foot keypoints (135 tota… | 23 | 34413 | maintenance |
| ok-oldking/ok-wuthering-waves ok-ww is an image-recognition-based automation tool for the game Wuthering Waves, supporting background operation, automatic combat, echo f… | 87 | 7167 | active |
| zxing/zxing ZXing ('Zebra Crossing') is an open-source, multi-format 1D/2D barcode image processing library implemented in Java, with ports to other la… | 73 | 34077 | maintenance |
| yangchris11/samurai SAMURAI is the official implementation of a zero-shot visual object tracker built on top of Segment Anything Model 2 (SAM 2), using a motio… | 27 | 7112 | active |
| chenfei-wu/TaskMatrix TaskMatrix (Visual ChatGPT) is a Python framework that connects ChatGPT with a suite of visual foundation models like Stable Diffusion, Gro… | 30 | 34003 | maintenance |
| OneDragon-Anything/ZenlessZoneZero-OneDragon A Python-based automation assistant for the game Zenless Zone Zero that uses image recognition and OCR to fully automate daily tasks, dunge… | 90 | 7029 | active |
| apple/corenet CoreNet is Apple's deep neural network training toolkit for training standard and novel small and large-scale models, including foundation … | 45 | 7007 | active |
| PaddlePaddle/models PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda… | 23 | 6932 | active |
| sczhou/ProPainter ProPainter is a PyTorch-based video inpainting model from ICCV 2023 that combines dual-domain propagation with a mask-guided sparse video T… | 22 | 6916 | stable |
| VAST-AI-Research/TripoSR TripoSR is an open-source model for fast feedforward 3D object reconstruction from a single image, developed by Tripo AI and Stability AI. … | 64 | 6888 | active |
| FoundationVision/ByteTrack ByteTrack is a PyTorch-based multi-object tracking (MOT) library implementing the ECCV 2022 paper 'Multi-Object Tracking by Associating Eve… | 32 | 6654 | stable |
| Yuliang-Liu/MonkeyOCR MonkeyOCR is a lightweight large multimodal model (LMM) for document parsing that uses a Structure-Recognition-Relation triplet paradigm to… | 60 | 6635 | active |
| scikit-image/scikit-image scikit-image is a Python library providing a collection of peer-reviewed image processing algorithms built on NumPy and SciPy. It offers ro… | 82 | 6577 | stable |
| iperov/DeepFaceLive DeepFaceLive is a real-time face-swap application for PC streaming and video calls, using trained face models (DFM) applied to webcam or vi… | 10 | 31011 | maintenance |
| openMVG/openMVG OpenMVG is a C++ library for multiple view geometry and Structure from Motion (SfM), providing end-to-end 3D reconstruction from images. It… | 48 | 6542 | active |
| AILab-CVC/YOLO-World YOLO-World is a real-time open-vocabulary object detection model and Python toolkit from Tencent AI Lab and HUST, published at CVPR 2024. I… | 29 | 6529 | active |
| open-mmlab/mmdetection3d MMDetection3D is OpenMMLab's next-generation platform for general 3D object detection, built on PyTorch. It provides a modular toolbox with… | 23 | 6518 | active |
| open-mmlab/mmcv MMCV is the foundational computer vision library for the OpenMMLab ecosystem, providing image/video I/O, data transformations, and CUDA ope… | 52 | 6470 | stable |
| sz3/libcimbar libcimbar is an optimized C++ implementation of the cimbar (color icon matrix) high-density 2D barcode format for air-gapped data transfer … | 95 | 6426 | active |
| OpenDroneMap/ODM OpenDroneMap (ODM) is an open source command line toolkit that processes aerial drone, balloon, or kite imagery into classified point cloud… | 90 | 6417 | active |
| madmaze/pytesseract Python-tesseract is a Python wrapper for Google's Tesseract-OCR engine that recognizes and extracts text embedded in images. It supports al… | 64 | 6383 | stable |
| KevinMusgrave/pytorch-metric-learning A PyTorch library providing modular losses, miners, samplers, and testers for deep metric learning. It supports building complete train/tes… | 48 | 6339 | active |
| mindee/doctr docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe… | 90 | 6315 | active |
| szad670401/HyperLPR HyperLPR3 is a high-performance open-source framework for recognizing Chinese license plates, built with deep learning and available as a P… | 27 | 6255 | active |
| RangiLyu/nanodet NanoDet-Plus is a super fast, lightweight anchor-free object detection model implemented in PyTorch, with model sizes as small as 980KB (IN… | 23 | 6252 | stable |
| PaddlePaddle/PaddleX PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti… | 92 | 6251 | active |
| ByteDance-Seed/Depth-Anything-3 Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,… | 59 | 6213 | active |
| ByteDance-Seed/Bagel BAGEL is an open-source unified multimodal foundation model with 7B active parameters (14B total) trained on interleaved multimodal data. I… | 55 | 6159 | active |
| open-edge-platform/anomalib Anomalib is a Python deep learning library for anomaly detection, offering state-of-the-art unsupervised algorithms for detecting and local… | 98 | 6088 | active |
| shimat/opencvsharp OpenCvSharp is a cross-platform .NET wrapper for the OpenCV computer vision library, published as NuGet packages with bundled native binari… | 98 | 6072 | active |
| bytedance/LatentSync LatentSync is an end-to-end lip-sync framework from ByteDance based on audio-conditioned latent diffusion models, using Stable Diffusion to… | 33 | 6026 | active |
| CGAL/cgal CGAL (Computational Geometry Algorithms Library) is a mature C++ library providing efficient and reliable algorithms for computational geom… | 95 | 6024 | stable |
| om-ai-lab/VLM-R1 VLM-R1 is a framework for training R1-style large vision-language models using reinforcement learning (GRPO) on top of Qwen2.5-VL. It provi… | 63 | 6015 | active |
| Doubiiu/ToonCrafter ToonCrafter is a generative model that interpolates two cartoon images into a short animation by leveraging pre-trained image-to-video diff… | 29 | 6003 | stable |
| OFA-Sys/Chinese-CLIP Chinese-CLIP is a Chinese version of the CLIP model trained on ~200 million Chinese image-text pairs, built on open_clip. It provides APIs,… | 66 | 5998 | active |
| journeyapps/zxing-android-embedded An Android barcode scanning library built on the ZXing decoder, usable via Intents or embedded in an Activity for custom UI. It supports po… | 23 | 5931 | stable |
| ChaoningZhang/MobileSAM MobileSAM is the official implementation of a lightweight version of Meta's Segment Anything Model (SAM), replacing the heavyweight image e… | 65 | 5858 | stable |
| PaddlePaddle/PaddleClas PaddleClas is a Python library and toolkit for image classification, recognition, and retrieval built on the PaddlePaddle deep learning fra… | 66 | 5838 | active |
| cnr-isti-vclab/meshlab MeshLab is an open source, portable, and extensible system for processing and editing unstructured large 3D triangular meshes, built on the… | 67 | 5802 | active |
| pjreddie/darknet Darknet is an open-source neural network framework written in C and CUDA, best known as the original home of the YOLO real-time object dete… | 32 | 26492 | maintenance |
| DeepLabCut/DeepLabCut DeepLabCut is an open-source Python toolbox for markerless 2D and 3D pose estimation of user-defined body parts using deep neural networks … | 89 | 5745 | stable |
| NVIDIA/DALI NVIDIA DALI is a GPU-accelerated data loading and preprocessing library with optimized building blocks and an execution engine for deep lea… | 92 | 5734 | active |
| ladaapp/lada Lada is an open-source tool with both GUI and CLI that restores pixelated/mosaic regions in videos, primarily targeting JAV (Japanese adult… | 65 | 5702 | active |
| open-mmlab/OpenPCDet OpenPCDet is a PyTorch-based open-source toolbox for LiDAR-based 3D object detection. It provides official implementations of models like P… | 53 | 5692 | active |
| apple/ml-depth-pro Depth Pro is Apple's reference implementation of a foundation model for zero-shot metric monocular depth estimation, producing sharp high-r… | 30 | 5683 | active |
| idealo/imagededup imagededup is a Python library for finding exact and near-duplicate images in a collection using perceptual hashing algorithms (PHash, DHas… | 48 | 5666 | stable |
| Fanghua-Yu/SUPIR SUPIR is a Python-based photo-realistic image restoration system built on SDXL diffusion priors and LLaVA captioning, presented at CVPR 202… | 36 | 5649 | active |
| Vision-CAIR/MiniGPT-4 Official code for MiniGPT-4 and MiniGPT-v2, vision-language models that align a frozen visual encoder with a frozen LLM (Vicuna) via a sing… | 30 | 25627 | maintenance |
| nerfstudio-project/gsplat gsplat is an open-source Python library with CUDA-accelerated, differentiable rasterization of Gaussians, based on 3D Gaussian Splatting fo… | 70 | 5589 | active |
| matterport/Mask_RCNN A Python implementation of the Mask R-CNN model for object detection and instance segmentation, built on Keras and TensorFlow with a ResNet… | 23 | 25567 | maintenance |
| ngxson/smolvlm-realtime-webcam A browser-based demo that streams webcam frames to a llama.cpp server running SmolVLM 500M for real-time object detection and scene descrip… | 29 | 5571 | active |
| lyuwenyu/RT-DETR Official implementation of RT-DETR and RT-DETRv2, real-time object detection transformers that outperform YOLO models, in PyTorch and Paddl… | 74 | 5476 | active |
| obss/sahi SAHI (Slicing Aided Hyper Inference) is a Python vision library for detecting small objects in large images via sliced/tiled inference, wor… | 99 | 5473 | active |