domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| shitagaki-lab/see-through A research framework from a SIGGRAPH 2026 paper that decomposes a single anime character illustration into up to 23 fully inpainted, semant… | 58 | 3641 | active |
| cmusatyalab/openface OpenFace is a free and open source Python and Torch implementation of face recognition based on Google's FaceNet deep neural network. It ge… | 65 | 15438 | maintenance |
| facebookresearch/detr DETR is Facebook Research's PyTorch implementation of Detection Transformer, an end-to-end object detection model that replaces hand-crafte… | 10 | 15354 | maintenance |
| albumentations-team/albumentations Albumentations is a fast, flexible Python image augmentation library for computer vision, supporting images, masks, bounding boxes, keypoin… | 10 | 15315 | maintenance |
| MrNeRF/LichtFeld-Studio LichtFeld Studio is a native open-source desktop application for 3D Gaussian Splatting that combines training, real-time inspection, splat … | 92 | 3594 | active |
| ZhaoJ9014/face.evoLVe A high-performance face recognition library built on PaddlePaddle and PyTorch, providing comprehensive tools for face-related analytics and… | 38 | 3589 | active |
| AiuniAI/Unique3D Unique3D is the official implementation of a NeurIPS 2024 paper that generates high-quality textured 3D meshes from a single image in about… | 38 | 3579 | active |
| SkalskiP/make-sense makesense.ai is a free, browser-based tool for labeling photos to prepare datasets for computer vision projects. It runs entirely client-si… | 23 | 3562 | active |
| GVCLab/PersonaLive PersonaLive is a diffusion-based framework for real-time, streamable portrait image animation, generating infinite-length expressive talkin… | 53 | 3552 | active |
| ToTheBeginning/PuLID PuLID is the official PyTorch implementation of a NeurIPS 2024 method for inserting a specific person's identity into text-to-image generat… | 40 | 3550 | active |
| AliaksandrSiarohin/first-order-model Official PyTorch/Jupyter implementation of the First Order Motion Model for image animation (NeurIPS 2019). It animates a static source ima… | 32 | 15015 | maintenance |
| sparkjsdev/spark Spark is an advanced 3D Gaussian Splatting renderer library for THREE.js, built in TypeScript by World Labs. It integrates with the THREE.j… | 81 | 3529 | active |
| google-research/big_vision Google Research's official Jax/Flax codebase for training large-scale vision models such as Vision Transformer, SigLIP, MLP-Mixer, and LiT … | 42 | 3528 | active |
| cszn/KAIR A PyTorch image restoration toolbox providing training and testing code for many restoration models including DnCNN, FFDNet, SRMD, USRNet, … | 23 | 3523 | active |
| MashiroSaber03/Saber-Translator Saber-Translator is an AI-powered manga translation application that detects speech bubbles, OCRs Japanese text, translates it, inpaints th… | 78 | 3516 | active |
| NVlabs/FoundationPose FoundationPose is NVIDIA's unified foundation model for 6D object pose estimation and tracking of novel objects, supporting both model-base… | 62 | 3516 | active |
| MooreThreads/Moore-AnimateAnyone An open-source reproduction of AnimateAnyone that animates a character from a single reference image using pose sequences from a driving vi… | 26 | 3514 | active |
| visionml/pytracking PyTracking is a PyTorch-based framework for visual object tracking and video object segmentation, providing official implementations of tra… | 23 | 3514 | active |
| PKU-YuanGroup/Video-LLaVA Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into… | 27 | 3500 | active |
| Anttwo/SuGaR SuGaR is the official PyTorch implementation of a CVPR 2024 method that extracts accurate, editable meshes from 3D Gaussian Splatting recon… | 27 | 3495 | active |
| aleju/imgaug imgaug is a Python library for augmenting images in machine learning experiments, converting a small set of input images into a much larger… | 23 | 14741 | maintenance |
| facebookresearch/ijepa Official PyTorch implementation of I-JEPA, a self-supervised learning method that predicts latent representations of image regions from oth… | 10 | 3489 | active |
| AliceVision Meshroom is an open-source, node-based visual programming application for building and executing data processing pipelines, best known for … | 69 | 3486 | active |
| vietanhdev/anylabeling AnyLabeling is a desktop image annotation tool that combines LabelImg/Labelme-style manual labeling with AI-assisted auto-labeling. It runs… | 85 | 3463 | active |
| NVlabs/Eagle Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L… | 64 | 3462 | active |
| RainerKuemmerle/g2o g2o is an open-source C++ framework for optimizing graph-based nonlinear error functions, commonly used for nonlinear least squares problem… | 67 | 3461 | stable |
| facebookresearch/sam-3d-body SAM 3D Body is a promptable model for single-image full-body 3D human mesh recovery (HMR), estimating body, feet, and hand pose using the M… | 48 | 3461 | active |
| nihui/waifu2x-ncnn-vulkan A portable command-line tool implementing the waifu2x anime-style image upscaler and denoiser using the ncnn inference framework with the V… | 63 | 3456 | active |
| roflcoopter/viseron Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti… | 98 | 3431 | active |
| davidsandberg/facenet A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC… | 32 | 14343 | maintenance |
| luigifreda/pyslam pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras with a wide range of local and global fe… | 77 | 3401 | active |
| IQA-PyTorch A pure Python/PyTorch toolbox for image quality assessment (IQA) providing GPU-accelerated reimplementations of many full-reference and no-… | 82 | 3380 | active |
| deepseek-ai/DeepSeek-OCR-2 DeepSeek-OCR 2 is an open-source vision-language model and inference toolkit implementing 'Visual Causal Flow' for optical character recogn… | 44 | 3379 | active |
| WongKinYiu/yolov7 Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques… | 23 | 14139 | maintenance |
| jenly1314/ZXingLite ZXingLite is a streamlined, fast Android library built on ZXing for scanning and generating QR codes and barcodes, with fully customizable … | 82 | 3367 | active |
| OpenTalker/SadTalker SadTalker is a CVPR 2023 deep learning tool that generates realistic talking head videos from a single portrait image and an audio clip by … | 22 | 14040 | maintenance |
| VainF/Torch-Pruning Torch-Pruning is a PyTorch framework for structural neural network pruning based on the DepGraph algorithm from CVPR 2023. It automatically… | 50 | 3348 | active |
| OpenGVLab/Ask-Anything VideoChat/Ask-Anything is a family of multimodal chat models and demos that combine video understanding with large language models, letting… | 72 | 3346 | active |
| nihui/opencv-mobile opencv-mobile provides minimal, prebuilt OpenCV binary packages for Android, iOS, ARM Linux, Windows, Linux, macOS, HarmonyOS, WebAssembly,… | 85 | 3345 | active |
| RKNN-Toolkit2 RKNN-Toolkit2 is Rockchip's SDK for converting trained neural network models into RKNN format and deploying them on Rockchip NPU chips like… | 36 | 3313 | active |
| Peterande/D-FINE D-FINE is the official PyTorch implementation of an ICLR 2025 Spotlight paper that redefines the regression task in DETR-style detectors as… | 67 | 3305 | active |
| XiaoMi/xiaomi-miloco Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an… | 82 | 3292 | active |
| hbb1/2d-gaussian-splatting Official implementation of 2D Gaussian Splatting (2DGS), a SIGGRAPH 2024 method that represents scenes as 2D oriented Gaussian disks for ge… | 69 | 3279 | stable |
| jixiaozhong/Sonic Sonic is the official PyTorch implementation of the CVPR 2025 paper 'Sonic: Shifting Focus to Global Audio Perception in Portrait Animation… | 49 | 3273 | active |
| vladmandic/human Human is a JavaScript/TypeScript library built on TensorFlow.js that combines multiple ML models for 3D face detection and recognition, bod… | 48 | 3264 | active |
| deepdoctection/deepdoctection deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c… | 98 | 3248 | active |
| Beckschen/TransUNet Official PyTorch implementation of TransUNet, a U-Net-style architecture that uses a Vision Transformer encoder for medical image segmentat… | 63 | 3234 | stable |
| mit-han-lab/bevfusion BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re… | 10 | 3230 | stable |
| Jittor/jittor Jittor is a high-performance deep learning framework from Tsinghua University based on just-in-time (JIT) compilation and meta-operators, w… | 67 | 3229 | active |
| breezedeus/Pix2Text Pix2Text is an open-source Python tool that recognizes layouts, tables, math formulas (LaTeX), and text in images and converts them into Ma… | 99 | 3227 | active |
| MzeroMiko/VMamba VMamba is a PyTorch implementation of a visual state space model (SSM) vision backbone based on Mamba, featuring 2D Selective Scan (SS2D) f… | 21 | 3219 | active |
| kerlomz/captcha_trainer A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren… | 55 | 3213 | active |
| OpenGVLab/InternGPT InternGPT (iGPT) is an open-source demo platform for showcasing AI models through a pointing-language-driven visual interactive system, sup… | 29 | 3205 | active |
| VTK VTK (Visualization Toolkit) is an open-source C++ library for 3D graphics, image processing, volume rendering, and scientific visualization… | 77 | 3201 | stable |
| PJLab-ADG/SensorsCalibration OpenCalib is a C++ multi-sensor calibration toolbox for autonomous driving that calibrates IMU, LiDAR, camera, and radar sensors, both intr… | 32 | 3200 | active |
| xianfei/SysMocap SysMocap is a cross-platform, video-driven real-time motion capture system that animates 3D virtual characters from webcam footage. It rend… | 78 | 3199 | active |
| prs-eth/Marigold Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus… | 52 | 3198 | active |
| Pointcept Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan… | 76 | 3196 | active |
| facebookresearch/dinov2 PyTorch implementation and pretrained models for DINOv2, a self-supervised vision transformer method from Meta AI that learns robust visual… | 68 | 13266 | maintenance |
| GuidoBartoli/sherloq Sherloq is an open-source digital image forensic toolset providing an integrated GUI environment for analyzing images for tampering and aut… | 74 | 3193 | active |
| ARM-software/ComputeLibrary Arm's Compute Library is a C++ collection of over 100 low-level machine learning and computer vision functions optimized for Arm Cortex-A/N… | 96 | 3183 | active |
| Rudrabha/Wav2Lip Wav2Lip is the official research code for the ACM Multimedia 2020 paper 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In … | 45 | 13182 | maintenance |
| rmurai0610/MASt3R-SLAM MASt3R-SLAM is a real-time monocular dense SLAM system built on the MASt3R two-view 3D reconstruction prior, producing globally consistent … | 43 | 3167 | active |
| TurixAI/TuriX-CUA TuriX is an open-source computer-use agent (CUA) that lets AI models take real actions on a desktop GUI - clicking, typing, and navigating … | 70 | 3156 | active |
| cleardusk/3DDFA_V2 3DDFA_V2 is the official PyTorch implementation of the ECCV 2020 paper 'Towards Fast, Accurate and Stable 3D Dense Face Alignment'. It regr… | 23 | 3149 | stable |
| megvii-research/NAFNet NAFNet is the official PyTorch implementation of a state-of-the-art image restoration network that removes nonlinear activation functions. … | 32 | 3148 | stable |
| automeris-io/WebPlotDigitizer WebPlotDigitizer is a computer vision assisted web application that extracts numerical data from images of charts and plots. It has been wi… | 65 | 3146 | active |
| z-x-yang/Segment-and-Track-Anything An open-source pipeline (SAM-Track) that segments and tracks arbitrary objects in videos using the Segment Anything Model for key-frame seg… | 61 | 3134 | active |
| otiai10/gosseract gosseract is a Go package that provides OCR (Optical Character Recognition) by binding to the Tesseract C++ library via cgo. It lets Go app… | 51 | 3130 | active |
| junyanz/CycleGAN A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I… | 32 | 12870 | maintenance |
| jina-ai/clip-as-service CLIP-as-service is a low-latency, high-scalability server for embedding images and text into fixed-length vectors using OpenAI's CLIP model… | 23 | 12836 | maintenance |
| naver/mast3r MASt3R is the official PyTorch implementation of 'Grounding Image Matching in 3D with MASt3R' (ECCV 2024), a model that performs dense 3D r… | 37 | 3088 | active |
| zxingify/zxingify-objc ZXingObjC is a full Objective-C port of the ZXing barcode image processing library, supporting encoding and decoding of many 1D and 2D barc… | 23 | 3075 | active |
| antimatter15/splat A WebGL-based real-time viewer for 3D Gaussian Splatting scenes, rendering photorealistic navigable 3D environments from photo-derived spla… | 51 | 3065 | active |
| AnyListen/tools-ocr Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model… | 23 | 3065 | active |
| rpng/open_vins OpenVINS is an open-source C++ platform for visual-inertial navigation research, centered on a filter-based (MSCKF/EKF) estimator that fuse… | 47 | 3057 | active |
| yuyuyzl/EasyVtuber EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int… | 62 | 3051 | active |
| thiagoalessio/tesseract-ocr-for-php A PHP wrapper library around the Tesseract OCR command-line binary, providing a fluent API for extracting text from images. It supports mul… | 61 | 3040 | stable |
| SharpAI/DeepCamera DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r… | 86 | 3019 | active |
| osmr/imgclsmob A research sandbox providing (re)implementations of numerous deep learning computer vision models for classification, segmentation, detecti… | 23 | 3016 | active |
| ZQPei/deep_sort_pytorch A PyTorch implementation of the Deep SORT multi-object tracking algorithm, pairing YOLOv3/YOLOv5 (or Mask R-CNN) detectors with a CNN re-id… | 32 | 3012 | active |
| Doubiiu/DynamiCrafter DynamiCrafter is an open-source research model that animates open-domain still images into short videos using pre-trained video diffusion p… | 27 | 3007 | active |
| williamyang1991/Rerender_A_Video The official PyTorch implementation of 'Rerender A Video', a SIGGRAPH Asia 2023 zero-shot text-guided video-to-video translation framework.… | 29 | 2999 | stable |
| MeiGen-AI/MultiTalk MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a … | 56 | 2992 | active |
| iscyy/ultralyticsPro A PyTorch-based collection of improved YOLO-family object detection models (YOLOv5 through YOLOv13, RT-DETR) with pluggable modules for bac… | 48 | 2954 | active |
| zju3dv/LoFTR LoFTR is a detector-free local image feature matching method using Transformers, released with PyTorch inference and training code plus pre… | 32 | 2950 | stable |
| sunsmarterjie/yolov12 YOLOv12 is a PyTorch implementation of attention-centric real-time object detectors, published at NeurIPS 2025. It provides detection model… | 59 | 2947 | active |
| microsoft/table-transformer Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un… | 23 | 2939 | active |
| KichangKim/DeepDanbooru DeepDanbooru is a Python/TensorFlow system that estimates Danbooru-style tags for anime-style girl images using multi-label classification.… | 63 | 2937 | active |
| sherlockchou86/VideoPipe VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates … | 54 | 2931 | active |
| jeeliz/jeelizFaceFilter A lightweight JavaScript/WebGL library for real-time face detection and tracking from a camera feed via WebRTC, designed for building augme… | 46 | 2928 | active |
| InternLM/InternLM-XComposer InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u… | 38 | 2925 | active |
| xelatihy/yocto-gl Yocto/GL is a collection of small C++17 libraries for building physically-based graphics algorithms, written in a data-oriented style and r… | 23 | 2925 | active |
| ogkalu2/comic-translate An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language… | 90 | 2911 | active |
| mitsuba-renderer/mitsuba3 Mitsuba 3 is a research-oriented, retargetable rendering system for forward and inverse light transport simulation, written in C++17 on top… | 94 | 2899 | active |
| learnables/learn2learn learn2learn is a PyTorch library for meta-learning research, providing utilities for few-shot task creation, high-level wrappers for algori… | 48 | 2893 | active |
| nimiq/qr-scanner A lightweight JavaScript/TypeScript QR code scanner library based on Cosmo Wolfe's port of Google's ZXing library. It supports webcam video… | 32 | 2886 | stable |
| NVlabs/FoundationStereo FoundationStereo is NVIDIA's official PyTorch implementation of a foundation model for zero-shot stereo depth estimation, published as a CV… | 47 | 2874 | active |
| UX-Decoder/Semantic-SAM Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired … | 33 | 2854 | active |
| openmv/openmv OpenMV is an open-source machine vision platform consisting of camera hardware firmware programmable in Python 3 (MicroPython). The firmwar… | 91 | 2850 | active |