domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| hacksider/Deep-Live-Cam Deep-Live-Cam is a Python application that performs real-time face swapping on webcam feeds and one-click video deepfakes using only a sing… | 89 | 96140 | active |
| ruvnet/RuView RuView is a WiFi sensing platform that uses Channel State Information (CSI) from commodity WiFi hardware like ESP32 to detect presence, tra… | 81 | 91756 | active |
| OpenCV OpenCV is the de facto open-source computer vision library, providing thousands of optimized algorithms for image and video processing, fea… | 89 | 90613 | stable |
| PaddlePaddle/PaddleOCR PaddleOCR is a multilingual OCR and document parsing toolkit built on PaddlePaddle that converts images and PDFs into structured data like … | 93 | 88312 | stable |
| tensorflow/models The TensorFlow Model Garden is a repository of official and community implementations of state-of-the-art machine learning models built wit… | 85 | 77652 | active |
| Tesseract OCR Tesseract is an open-source OCR engine consisting of the libtesseract library and a command-line program, using an LSTM-based neural networ… | 86 | 76200 | stable |
| ultralytics/ultralytics Ultralytics YOLO is a Python package and CLI providing a family of real-time computer vision models (YOLO26, YOLO11, YOLOv8) for object det… | 95 | 60991 | active |
| facebookresearch/segment-anything Segment Anything Model (SAM) from Meta AI is a promptable image segmentation foundation model that produces high-quality object masks from … | 30 | 54759 | stable |
| roboflow/supervision Supervision is a Python library of reusable computer vision tools that bridges the gap between detection/segmentation/classification models… | 95 | 49745 | active |
| hiroi-sora/Umi-OCR Umi-OCR is a free, open-source, fully offline OCR application for Windows and Linux with a Qt/QML GUI. It supports screenshot OCR, batch im… | 47 | 46882 | stable |
| naptha/tesseract.js Tesseract.js is a pure JavaScript port of the Tesseract OCR engine that extracts text from images in over 100 languages. It runs in the bro… | 70 | 38671 | active |
| huggingface/pytorch-image-models PyTorch Image Models (timm) is a Python library offering the largest collection of PyTorch image encoder/backbone architectures with 700+ p… | 93 | 37099 | active |
| google-ai-edge/mediapipe MediaPipe is Google's cross-platform framework for deploying on-device machine learning solutions for live and streaming media. It provides… | 94 | 36731 | stable |
| Real-ESRGAN Real-ESRGAN is a deep learning project for practical image and video restoration via super-resolution, with pretrained models for photos an… | 23 | 36593 | stable |
| XingangPan/DragGAN Official PyTorch implementation of DragGAN (SIGGRAPH 2023), an interactive point-based image manipulation method built on StyleGAN3. Users … | 29 | 35755 | stable |
| Frigate Frigate is an open-source, self-hosted network video recorder (NVR) that performs real-time AI object detection on IP camera feeds locally … | 92 | 35407 | active |
| facebookresearch/detectron2 Detectron2 is Facebook AI Research's PyTorch-based library for state-of-the-art object detection, instance/panoptic segmentation, and other… | 67 | 34688 | stable |
| openai/CLIP OpenAI's CLIP is a PyTorch library providing pretrained contrastive language-image models that encode images and text into a shared embeddi… | 66 | 34236 | stable |
| FreeCAD/FreeCAD FreeCAD is a free and open-source, cross-platform 3D parametric CAD modeler for designing real-life objects of any size, built on the OpenC… | 89 | 33083 | stable |
| open-mmlab/mmdetection MMDetection is OpenMMLab's PyTorch-based toolbox and benchmark for object detection and instance/panoptic segmentation. It provides a large… | 23 | 32892 | stable |
| JaidedAI/EasyOCR EasyOCR is a ready-to-use Python OCR library built on PyTorch that extracts text from images, supporting 80+ languages and popular writing … | 48 | 29942 | stable |
| deepinsight/insightface InsightFace is an open-source 2D and 3D face analysis project providing state-of-the-art face detection, recognition, alignment, and face s… | 65 | 29580 | active |
| Label Studio Label Studio is an open-source data labeling and annotation platform supporting images, audio, text, video, and time series with a web UI a… | 86 | 28150 | active |
| OpenBMB/MiniCPM-V MiniCPM-V and MiniCPM-o are a series of small multimodal large language models for efficient image, video, and audio understanding, deploya… | 61 | 26240 | active |
| lucidrains/vit-pytorch A PyTorch library implementing the Vision Transformer (ViT) and dozens of ViT variants (NaViT, MaxViT, MobileViT, Dino, masked autoencoders… | 87 | 25488 | active |
| microsoft/OmniParser OmniParser is a screen parsing tool from Microsoft that converts UI screenshots into structured, understandable elements to ground vision-l… | 60 | 25310 | active |
| junyanz/pytorch-CycleGAN-and-pix2pix Official PyTorch implementations of CycleGAN and pix2pix for paired and unpaired image-to-image translation. It includes training and testi… | 48 | 25232 | stable |
| haotian-liu/LLaVA LLaVA (Large Language and Vision Assistant) is an open-source multimodal large language model framework implementing visual instruction tun… | 20 | 25000 | active |
| baidu/Unlimited-OCR Baidu's Unlimited-OCR is an open vision-language OCR model for one-shot long-horizon document parsing, extending DeepSeek-OCR. It provides … | 56 | 24569 | active |
| danielgatis/rembg Rembg is a Python tool for removing image backgrounds using U2Net-based deep learning models. It can be used as a CLI, Python library, HTTP… | 98 | 24449 | active |
| graphdeco-inria/gaussian-splatting The official reference implementation of 3D Gaussian Splatting, a method for real-time radiance field rendering that reconstructs scenes fr… | 50 | 23425 | active |
| serengil/deepface DeepFace is a lightweight Python library for face recognition and facial attribute analysis, wrapping state-of-the-art models like VGG-Face… | 89 | 23340 | stable |
| MAA (MaaAssistantArknights) MAA (MAA Assistant Arknights) is a C++ desktop assistant for the mobile game Arknights that automates daily tasks using image recognition. … | 94 | 22796 | active |
| huggingface/datasets Hugging Face Datasets is a Python library providing one-line access to hundreds of thousands of public datasets on the Hugging Face Hub acr… | 98 | 21870 | stable |
| Zeyi-Lin/HivisionIDPhotos HivisionIDPhotos is a lightweight AI tool that generates standard ID/passport photos from user images using offline matting models that run… | 71 | 21420 | active |
| datalab-to/surya Surya is a 650M parameter OCR toolkit from Datalab providing state-of-the-art text recognition, layout analysis, reading order detection, a… | 86 | 21318 | active |
| bloc97/Anime4K Anime4K is a set of open-source, high-quality real-time anime upscaling and denoising algorithms implemented as GLSL shaders, primarily for… | 23 | 21295 | stable |
| QwenLM/Qwen3-VL Qwen3-VL is a series of open-weight multimodal vision-language models from Alibaba's Qwen team, available in Dense and MoE architectures wi… | 52 | 19847 | active |
| facebookresearch/sam2 Official code for Meta's Segment Anything Model 2 (SAM 2), a foundation model for promptable visual segmentation in images and videos. It i… | 61 | 19770 | active |
| KlingAIResearch/LivePortrait LivePortrait is a Python-based portrait animation tool from Kuaishou Technology that synthesizes lifelike videos from a single source image… | 62 | 18969 | active |
| sczhou/CodeFormer CodeFormer is a PyTorch-based blind face restoration model using a codebook lookup transformer, published at NeurIPS 2022. It restores and … | 46 | 18117 | stable |
| pytorch/vision torchvision is the official PyTorch companion library providing datasets, model architectures, and image/video transformations for computer… | 93 | 17885 | stable |
| deepseek-ai/Janus Janus-Series is DeepSeek's family of unified multimodal models (Janus, Janus-Pro, JanusFlow) that combine multimodal understanding and imag… | 24 | 17757 | active |
| IDEA-Research/Grounded-Segment-Anything Grounded-Segment-Anything (Grounded SAM) combines Grounding DINO with Segment Anything to detect and segment arbitrary objects from text pr… | 30 | 17710 | active |
| NVlabs/instant-ngp NVIDIA's implementation of instant neural graphics primitives, training NeRFs, signed distance functions, neural images, and neural volumes… | 52 | 17535 | stable |
| Robbyant/lingbot-map LingBot-Map is a feed-forward 3D foundation model that reconstructs scenes from streaming image data using a Geometric Context Transformer.… | 58 | 16705 | active |
| cvat-ai/cvat CVAT (Computer Vision Annotation Tool) is an open-source, self-hosted platform for annotating images, videos, and 3D point clouds to build … | 95 | 16600 | active |
| lukas-blecher/LaTeX-OCR pix2tex (LaTeX-OCR) is a PyTorch-based vision transformer model that converts images of math formulas into LaTeX code. It ships as a pip-in… | 24 | 16547 | stable |
| huggingface/transformers.js Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) … | 91 | 16270 | active |
| wkentaro/labelme Labelme is a graphical image annotation tool written in Python with a Qt interface, supporting polygon, rectangle, oriented rectangle, circ… | 99 | 16130 | active |
| microsoft/Swin-Transformer Official PyTorch implementation of the Swin Transformer, a hierarchical vision transformer using shifted windows that serves as a general-p… | 32 | 16051 | stable |
| babalae/better-genshin-impact BetterGI is a free, open-source Windows desktop application that automates gameplay in Genshin Impact using computer vision, OCR, and YOLO-… | 95 | 15067 | active |
| duixcom/Duix-Avatar Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea… | 57 | 14871 | active |
| tensorflow/tfjs-models A collection of pre-trained machine learning models ported to TensorFlow.js, published as npm packages for use in JavaScript projects. Mode… | 63 | 14793 | active |
| HumanAIGC/AnimateAnyone Animate Anyone is the official research implementation of a diffusion-based image-to-video synthesis method that animates a static characte… | 46 | 14791 | active |
| ddddocr DdddOcr is a Python library for offline, local recognition of various CAPTCHA types, including alphanumeric, Chinese character, and slider … | 64 | 14665 | active |
| dlib Dlib is a modern C++ toolkit containing machine learning algorithms, deep learning tools, computer vision, linear algebra, and general-purp… | 86 | 14431 | stable |
| PaddlePaddle/PaddleDetection PaddleDetection is an object detection toolkit built on the PaddlePaddle deep learning framework. It provides implementations of detection,… | 73 | 14389 | active |
| facebookresearch/vggt VGGT (Visual Geometry Grounded Transformer) is a feed-forward transformer model from Meta AI and Oxford VGG that infers 3D geometry—camera … | 58 | 14292 | active |
| mlfoundations/open_clip OpenCLIP is an open-source PyTorch implementation of CLIP and related multimodal contrastive models, with many pretrained image/text checkp… | 86 | 14095 | active |
| img2threejs/img2threejs A tool that reconstructs objects from reference images as code-only, procedural Three.js models rather than meshes or photogrammetry. It pr… | 80 | 14018 | active |
| Open3D Open3D is an open-source C++ and Python library for 3D data processing, offering data structures, algorithms, and pipelines for point cloud… | 67 | 13913 | active |
| jacobgil/pytorch-grad-cam A PyTorch library providing state-of-the-art pixel attribution (saliency) methods like GradCAM, ScoreCAM, and AblationCAM for explainable A… | 76 | 12958 | active |
| jwagner/smartcrop.js smartcrop.js is a JavaScript library that implements a content-aware algorithm to find good crops for images. It runs in the browser, in No… | 23 | 12955 | stable |
| ShiqiYu/libfacedetection An open-source C++ library for CNN-based face detection in images, with the model embedded as static C source so it has no external depende… | 63 | 12784 | stable |
| google-research/vision_transformer Google Research's official JAX/Flax implementation of Vision Transformer (ViT) and MLP-Mixer architectures, with released pretrained checkp… | 75 | 12683 | stable |
| colmap/colmap COLMAP is a general-purpose Structure-from-Motion (SfM) and Multi-View Stereo (MVS) pipeline for reconstructing 3D models from ordered or u… | 98 | 12564 | active |
| YaoFANGUK/video-subtitle-remover An AI-based desktop application that removes hard-coded subtitles and text-like watermarks from videos and images using deep learning inpai… | 71 | 12553 | active |
| DayBreak-u/chineseocr_lite An ultra-lightweight Chinese OCR toolkit combining DBNet text detection, CRNN text recognition, and an angle classifier, with total model s… | 70 | 12339 | active |
| simular-ai/Agent-S Agent S is an open-source agentic framework that uses multimodal LLMs to operate computers like a human, controlling GUIs via clicking, typ… | 70 | 12193 | active |
| xmu-xiaoma666/External-Attention-pytorch A PyTorch library (fightingcv-attention) providing clean, minimal implementations of numerous attention mechanisms, MLP variants, re-parame… | 65 | 12183 | active |
| datalab-to/chandra Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO… | 71 | 12171 | active |
| instantX-research/InstantID InstantID is a tuning-free, zero-shot identity-preserving image generation method built on diffusion models, generating customized images i… | 26 | 11987 | active |
| nerfstudio-project/nerfstudio Nerfstudio is a Python library and CLI toolkit providing a simple, modular API for creating, training, and testing Neural Radiance Fields (… | 38 | 11934 | active |
| qubvel-org/segmentation_models.pytorch A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar… | 70 | 11706 | stable |
| milesial/Pytorch-UNet A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva… | 23 | 11613 | active |
| facebookresearch/sam3 Official code for Meta's Segment Anything Model 3 (SAM 3), a unified foundation model for promptable segmentation in images and videos. It … | 63 | 11487 | active |
| UI-TARS UI-TARS is ByteDance's open-source multimodal AI agent stack, comprising Agent TARS (a CLI/Web UI multimodal agent that operates terminals,… | 49 | 11389 | active |
| rerun-io/rerun Rerun is an open-source SDK and viewer for logging, storing, querying, and visualizing multi-rate multimodal data such as images, point clo… | 99 | 11362 | active |
| THU-MIG/yolov10 YOLOv10 is a real-time end-to-end object detection model family that removes NMS post-processing via consistent dual assignments and optimi… | 20 | 11336 | active |
| kornia/kornia Kornia is a differentiable computer vision library built on PyTorch, offering GPU-accelerated image processing, augmentations, and geometri… | 86 | 11327 | active |
| salesforce/LAVIS LAVIS is a Python library from Salesforce AI Research providing a unified toolkit for language-vision (multimodal) intelligence, including … | 61 | 11262 | active |
| facebookresearch/dinov3 Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t… | 59 | 11249 | active |
| ultralytics/yolov5 Ultralytics YOLOv5 is a PyTorch-based computer vision model family for real-time object detection, instance segmentation, and image classif… | 67 | 57929 | maintenance |
| PointCloudLibrary/pcl The Point Cloud Library (PCL) is a large-scale, modular open-source C++ library for 2D/3D image and point cloud processing. It provides sta… | 68 | 11101 | stable |
| voxel51/fiftyone FiftyOne is an open-source Python library and GUI app for building high-quality computer vision datasets and models. It enables visualizing… | 99 | 11042 | active |
| ageitgey/face_recognition A Python library and command-line tool providing a simple API for face detection, facial landmark extraction, and face recognition, built o… | 63 | 56684 | maintenance |
| microsoft/TRELLIS.2 TRELLIS.2 is a 4B-parameter state-of-the-art 3D generative model from Microsoft for high-fidelity image-to-3D asset generation. It uses a f… | 57 | 10869 | active |
| OpenVINO OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su… | 95 | 10740 | stable |
| autogluon/autogluon AutoGluon is an AutoML library that automates machine learning on tabular data, time series, text, and images with just a few lines of Pyth… | 90 | 10617 | active |
| Megvii-BaseDetection/YOLOX YOLOX is a high-performance anchor-free YOLO object detection model family implemented in PyTorch, with pretrained weights and export suppo… | 34 | 10587 | stable |
| IDEA-Research/GroundingDINO Official PyTorch implementation of Grounding DINO, an open-set object detector that fuses a Transformer-based DINO detector with grounded v… | 21 | 10515 | stable |
| esimov/caire Caire is a content-aware image resize library written in Go, based on the seam carving algorithm. It intelligently shrinks or enlarges imag… | 31 | 10465 | active |
| openframeworks/openFrameworks openFrameworks is an open-source C++ toolkit for creative coding that wraps common libraries like OpenGL, OpenCV, and audio/video libraries… | 87 | 10419 | stable |
| zyddnys/manga-image-translator A Python tool that automatically detects, OCRs, translates, inpaints, and re-typesets text in images, primarily for manga and comics. It ru… | 65 | 10345 | active |
| CVHub520/X-AnyLabeling X-AnyLabeling is a cross-platform desktop application for AI-assisted annotation of text, image, video, and multimodal data. It bundles bui… | 96 | 10212 | active |
| OpenGVLab/InternVL InternVL is a family of open-source multimodal large language models (vision-language models) that combine vision transformers with LLMs to… | 37 | 10146 | active |
| freemocap/freemocap FreeMoCap is a free, open-source, markerless motion capture system that uses ordinary cameras (webcams, GoPros, smartphones) to record and … | 98 | 10085 | active |
| m87-labs/moondream Moondream is an open-weight family of small, efficient vision language models (2B to 9B MoE) that perform image captioning, visual question… | 61 | 10014 | active |
| facebookresearch/pytorch3d PyTorch3D is Facebook AI Research's library of efficient, reusable components for deep learning with 3D data, built on PyTorch. It provides… | 74 | 9954 | active |
page 1 / 24 next →