function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| fangfufu/Linux-Fake-Background-Webcam A Python application that creates a virtual webcam on GNU/Linux with fake backgrounds, including background replacement, blurring, animated… | 63 | 1700 | active |
| kingsic/SGQRCode SGQRCode is an easy-to-use iOS library for scanning barcodes and QR codes, generating QR codes, and recognizing QR codes from images. It pr… | 23 | 1697 | stable |
| stephansturges/WALDO WALDO is an open-source object detection model based on a YOLOv8 backbone, trained with a synthetic data pipeline to detect people, vehicle… | 32 | 1695 | active |
| mebjas/html5-qrcode A lightweight, zero-dependency JavaScript/TypeScript library for scanning QR codes and barcodes in the browser using the device camera or l… | 48 | 6213 | maintenance |
| shubham-goel/4D-Humans 4DHumans is a Python research codebase implementing HMR 2.0, a transformer-based model for 3D human mesh recovery from single images, plus … | 59 | 1671 | active |
| ZrrSkywalker/Personalize-SAM PerSAM is the official implementation of 'Personalize Segment Anything Model with One Shot', which customizes the Segment Anything Model (S… | 29 | 1671 | active |
| nwojke/deep_sort A Python implementation of Deep SORT, a multi-object tracking algorithm that extends SORT with a deep appearance descriptor for robust data… | 36 | 6169 | maintenance |
| bytedance/Sa2VA Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understa… | 70 | 1666 | active |
| thtrieu/darkflow Darkflow is a Python library that translates Darknet's YOLO neural network definitions to TensorFlow, enabling real-time object detection a… | 32 | 6139 | maintenance |
| MgArcher/Text_select_captcha A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio… | 69 | 1656 | active |
| chineseocr A Python OCR toolkit that combines YOLO3-based text detection with CRNN/Dense recognition for Chinese and English text in natural scene ima… | 32 | 6123 | maintenance |
| InsightSoftwareConsortium/ITK The Insight Toolkit (ITK) is an open-source, cross-platform C++ library with Python bindings for image analysis, providing algorithms for p… | 94 | 1647 | stable |
| RQLuo/MixTeX-Latex-OCR MixTeX is a multimodal OCR application that recognizes LaTeX formulas, tables, and mixed Chinese/English text from images, running entirely… | 22 | 1637 | active |
| gaoxiang12/lightning-lm Lightning-LM is a C++ library providing a complete 3D LiDAR SLAM system with fast LIO front-end, real-time loop closure detection, and high… | 51 | 1632 | active |
| hku-mars/FAST-LIVO FAST-LIVO is a fast, tightly-coupled sparse-direct LiDAR-Inertial-Visual Odometry system combining a LIO subsystem that registers raw point… | 51 | 1631 | stable |
| HKUST-Aerial-Robotics/VINS-Mono VINS-Mono is a real-time SLAM framework for monocular visual-inertial systems, using an optimization-based sliding window formulation for h… | 32 | 6011 | maintenance |
| jeffffffli/HybrIK HybrIK is the official PyTorch implementation of a hybrid analytical-neural inverse kinematics method for 3D human pose and shape estimatio… | 23 | 1618 | stable |
| OminousIndustries/PhoneDriver PhoneDriver is a Python-based mobile automation agent that uses Qwen3-VL vision-language models to visually understand and control Android … | 38 | 1614 | active |
| Robbyant/lingbot-depth LingBot-Depth is a PyTorch-based model and toolkit for masked depth modeling that transforms incomplete, noisy depth sensor data into metri… | 56 | 1610 | active |
| meituan/YOLOv6 YOLOv6 is a single-stage object detection framework implemented in PyTorch, designed for industrial applications with a family of pretraine… | 23 | 5895 | maintenance |
| Kruk2/jasna Jasna is a GPU-accelerated tool that detects and restores mosaics in JAV videos and still images, with a native GUI, CLI, and streaming sup… | 81 | 1601 | active |
| autonomousvision/transfuser Official PyTorch implementation of TransFuser, a transformer-based multi-modal sensor fusion model for end-to-end autonomous driving, publi… | 53 | 1599 | stable |
| facebookresearch/fast3r Fast3R is the official PyTorch implementation of a CVPR 2025 model from Meta FAIR that reconstructs 3D scenes and estimates camera poses fr… | 10 | 1593 | active |
| yakhyo/uniface UniFace is a unified Python library for face analysis that bundles detection, recognition, landmark localization, face parsing, gaze estima… | 88 | 1586 | active |
| Core-Mate/OpenGUI OpenGUI is an Android GUI agent framework that lets AI agents see, understand, and operate real mobile app interfaces on physical Android d… | 77 | 1583 | active |
| ai-forever/ghost GHOST (Generative High-fidelity One Shot Transfer) is a one-shot face swap pipeline for images and videos, published as an IEEE paper and i… | 26 | 1582 | active |
| Layout-Parser/layout-parser LayoutParser is a Python toolkit for deep learning based document image analysis, offering unified APIs for layout detection models, layout… | 23 | 5774 | maintenance |
| Tencent/DepthCrafter DepthCrafter is a diffusion-based video depth estimation model from Tencent AI Lab that generates temporally consistent long depth sequence… | 38 | 1574 | active |
| Babyhamsta/Aimmy Aimmy is a universal AI-based aim alignment mechanism (aim assist) for gamers with impairments, built in C# using YOLOv8 models run via ONN… | 78 | 1568 | active |
| IDEA-Research/Rex-Omni Rex-Omni is a 3B-parameter multimodal large language model that unifies object detection, OCR, pointing, keypoint detection, and visual pro… | 47 | 1561 | active |
| real-stanford/universal_manipulation_interface Universal Manipulation Interface (UMI) is a data collection and policy learning framework that transfers in-the-wild human demonstrations i… | 63 | 1560 | active |
| ORB-HD/deface deface is a Python command-line tool that automatically anonymizes human faces in videos and photos. It detects faces in each frame and app… | 23 | 1558 | stable |
| yeemachine/kalidokit KalidoKit is a TypeScript library that converts 3D landmark outputs from Mediapipe/Tensorflow.js face, pose, and hand tracking models into … | 49 | 5699 | maintenance |
| ethz-asl/kalibr Kalibr is a visual-inertial calibration toolbox for camera systems and inertial measurement units. It supports multi-camera, camera-IMU, IM… | 32 | 5680 | maintenance |
| hustvl/MapTR MapTR is an end-to-end transformer-based framework for online vectorized HD map construction from camera imagery in autonomous driving. It … | 27 | 1542 | active |
| dindin0497/SeeIt SeeIt is an inclusive Android app with two accessibility modes: one that converts spoken speech into text and plays corresponding ASL (Amer… | 41 | 1540 | active |
| Arthur151/ROMP ROMP is a PyTorch-based library and pip-installable API (simple-romp) for real-time monocular multi-person 3D human mesh recovery, implemen… | 23 | 1538 | stable |
| studyhelperhelper/studyhelper An Android app disguised as a Sudoku game that automates earning daily points in the Xuexi Qiangguo (学习强国) app. It uses accessibility servi… | 23 | 1533 | active |
| JingyunLiang/SwinIR Official PyTorch implementation of SwinIR, a Swin Transformer-based model for image restoration tasks including super-resolution, denoising… | 23 | 5580 | maintenance |
| NirAharon/BoT-SORT BoT-SORT is a state-of-the-art multi-object tracker that combines motion and appearance information with camera motion compensation and an … | 32 | 1522 | active |
| Tencent/TFace TFace is a research platform from Tencent Youtu Lab for trusty face analysis, covering face recognition, face security (anti-spoofing), fac… | 57 | 1521 | active |
| caiyuanhao1998/Retinexformer Retinexformer is a one-stage Retinex-based Transformer model and toolbox for low-light image enhancement, published at ICCV 2023. It suppor… | 66 | 1518 | active |
| NVlabs/describe-anything Describe Anything Model (DAM) is a vision-language model that generates detailed descriptions of user-specified regions in images and video… | 32 | 1514 | active |
| BrokenSource/DepthFlow DepthFlow is a free, open-source Python application and library that converts still images into 3D parallax effect videos using monocular d… | 85 | 1512 | active |
| OneDragon-Anything/StarRailOneDragon StarRailOneDragon is a Python-based automation application for the game Honkai: Star Rail that handles daily tasks automatically, using scr… | 84 | 1512 | active |
| hkchengrex/Tracking-Anything-with-DEVA DEVA is a decoupled video segmentation framework that combines task-specific image-level segmentation models with a universal bi-directiona… | 27 | 1508 | stable |
| koide3/direct_visual_lidar_calibration A C++ toolbox for target-less, single-shot extrinsic calibration between LiDAR sensors and cameras, supporting spinning and non-repetitive … | 74 | 1507 | stable |
| czczup/ViT-Adapter Official PyTorch implementation of ViT-Adapter, an ICLR 2023 Spotlight paper introducing a pre-training-free adapter that lets plain Vision… | 34 | 1503 | stable |
| zapdos-labs/unblink Unblink is an AI-powered camera monitoring application that uses a vision language model (Qwen3-VL) to analyze camera frames, summarize act… | 50 | 1501 | active |
| meiqua/shape_based_matching A C++ library implementing Halcon-style shape-based matching (equivalent to LINE-MOD) using gradient orientation templates for robust 2D ob… | 32 | 1500 | active |
| introlab/rtabmap_ros RTAB-Map's ROS package providing real-time appearance-based SLAM (RGB-D, stereo, and LiDAR graph SLAM) as ROS 1 and ROS 2 nodes. It integra… | 76 | 1498 | active |
| ANTsX/ANTs Advanced Normalization Tools (ANTs) is a C++ command-line library for high-dimensional medical image registration and segmentation, built o… | 83 | 1497 | active |
| cheind/py-motmetrics py-motmetrics is a Python library for evaluating multiple object tracking (MOT) results with MOTChallenge-aligned CLEAR MOT, Identity, and … | 65 | 1487 | active |
| CUT3R/CUT3R CUT3R is the official PyTorch implementation of 'Continuous 3D Perception Model with Persistent State' (CVPR 2025 Oral), a stateful recurre… | 38 | 1486 | active |
| damiafuentes/DJITelloPy A Python library wrapping the official DJI Tello and Tello EDU SDKs, implementing all Tello commands including video streaming, state packe… | 24 | 1482 | active |
| hustvl/DiffusionDrive DiffusionDrive is a truncated diffusion model for real-time end-to-end autonomous driving, released as the official PyTorch implementation … | 44 | 1480 | active |
| pur1fying/blue_archive_auto_script BAAS (Blue Archive Auto Script) is a GUI-based automation program for the mobile game Blue Archive that runs against 16:9 emulator screens.… | 72 | 1476 | active |
| google/GNM GNM is an open ecosystem of parametric statistical human models and perception stacks from Google, starting with GNM Head, a high-fidelity … | 59 | 1469 | active |
| autonomousvision/mip-splatting Mip-Splatting is a research implementation of alias-free 3D Gaussian Splatting, introducing a 3D smoothing filter and 2D Mip filter to elim… | 27 | 1466 | active |
| dbolya/yolact YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++… | 51 | 5241 | maintenance |
| zylo117/Yet-Another-EfficientDet-Pytorch A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim… | 23 | 5238 | maintenance |
| NVIDIA-ISAAC-ROS/isaac_ros_visual_slam Isaac ROS Visual SLAM is a ROS 2 package providing GPU-accelerated visual simultaneous localization and mapping (VSLAM) using stereo visual… | 94 | 1448 | active |
| phonowell/genshin-impact-script A Genshin Impact automation script written in AutoHotkey that provides features like automatic fishing, item pickup, and dialogue skipping.… | 63 | 1446 | active |
| amdegroot/ssd.pytorch A PyTorch implementation of the Single Shot MultiBox Detector (SSD) object detection model from the 2016 paper by Wei Liu et al. It include… | 32 | 5221 | maintenance |
| serratus/quaggaJS QuaggaJS is a barcode-scanner library written entirely in JavaScript that supports real-time localization and decoding of barcode types suc… | 23 | 5207 | maintenance |
| sdcb/PaddleSharp A .NET/C# wrapper around Baidu's PaddleInference C API, providing PaddleOCR, PaddleDetection, rotation detection, Chinese segmentation, and… | 72 | 1441 | active |
| ThoughtfulDev/EagleEye EagleEye is a Python-based OSINT tool that identifies social media profiles (Instagram, Facebook, Twitter, YouTube) of a person using face … | 32 | 5197 | maintenance |
| valentinfrlch/ha-llmvision LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live … | 90 | 1440 | active |
| Walter0807/MotionBERT Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske… | 65 | 1439 | active |
| jakowenko/double-take Double Take is a self-hosted Docker application providing a unified UI and API for facial recognition. It abstracts multiple face detection… | 40 | 1436 | active |
| NVlabs/Fast-FoundationStereo Fast-FoundationStereo is NVIDIA's official PyTorch implementation of a real-time zero-shot stereo matching model family, accepted to CVPR 2… | 54 | 1432 | active |
| qupath/qupath QuPath is an open-source desktop application for bioimage analysis, aimed especially at digital pathology and whole-slide imaging. It provi… | 79 | 1430 | active |
| jasonmayes/Real-Time-Person-Removal A browser-based demo that removes people from complex video backgrounds in real time using TensorFlow.js. It learns the static background o… | 23 | 5154 | maintenance |
| natario1/CameraView CameraView is a well-documented, high-level Android library that simplifies capturing pictures and videos, wrapping Camera1 and Camera2 API… | 23 | 5125 | maintenance |
| ZheC/Realtime_Multi-Person_Pose_Estimation Reference implementation of the CVPR'17 paper 'Realtime Multi-Person Pose Estimation', a bottom-up approach that detects keypoints for mult… | 32 | 5123 | maintenance |
| facebookresearch/vggsfm VGGSfM is a deep learning-based Structure from Motion pipeline from Meta AI and Oxford VGG that recovers camera poses and 3D point clouds f… | 29 | 1421 | active |
| CSAILVision/semantic-segmentation-pytorch A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra… | 32 | 5078 | maintenance |
| khanamiryan/php-qrcode-detector-decoder A pure PHP library for detecting and decoding QR codes from images, ported from the ZXing library. It works without any PHP extensions beyo… | 37 | 1412 | active |
| NVlabs/BundleSDF BundleSDF is a CVPR 2023 research implementation from NVIDIA for near real-time 6-DoF pose tracking of unknown rigid objects from monocular… | 65 | 1410 | stable |
| IDEA-Research/DINO-X-API DINO-X API is a Python client library and examples for accessing DINO-X, a hosted unified vision model for open-world object detection and … | 36 | 1410 | active |
| OpenBMB/AgentCPM-GUI AgentCPM-GUI is an open-source 8B-parameter on-device GUI agent built on MiniCPM-V that takes Android screenshots as input and autonomously… | 46 | 1407 | active |
| imagej/imagej2 ImageJ2 is an open-source Java framework and application for processing and analyzing N-dimensional scientific image data, built on the Img… | 75 | 1399 | stable |
| nasa/astrobee NASA's open-source flight software for the Astrobee free-flying robots operating aboard the International Space Station, written primarily … | 48 | 1397 | active |
| yfeng95/PRNet PRNet is a Python/TensorFlow implementation of the ECCV 2018 Position Map Regression Network for joint 3D face reconstruction and dense ali… | 32 | 5013 | maintenance |
| om-ai-lab/OmDet OmDet-Turbo is a PyTorch implementation of a transformer-based open-vocabulary object detection model that detects arbitrary user-defined o… | 57 | 1393 | active |
| yipianfengye/android-zxingLibrary An Android library wrapping ZXing that lets developers integrate QR code and barcode scanning into their apps with just a few lines of code… | 32 | 4995 | maintenance |
| Junyi42/monst3r MonST3R is the official PyTorch implementation of an ICLR 2025 paper that estimates per-timestep geometry (pointmaps) from dynamic videos i… | 36 | 1386 | active |
| zhixuhao/unet A Keras implementation of the U-Net convolutional network architecture for image segmentation, based on the original biomedical segmentatio… | 66 | 4941 | maintenance |
| autonomousvision/unimatch UniMatch is a PyTorch research library implementing a unified transformer-based model for optical flow, stereo matching, and depth estimati… | 32 | 1379 | stable |
| yanx27/Pointnet_Pointnet2_pytorch A pure PyTorch implementation of the PointNet and PointNet++ deep learning architectures for point cloud processing. It includes training a… | 32 | 4936 | maintenance |
| siyuanliii/masa Official PyTorch implementation of MASA (CVPR 2024 Highlight), a universal instance appearance model that learns to match any objects acros… | 33 | 1377 | active |
| hustvl/VAD VAD is an end-to-end autonomous driving framework that models the driving scene as a fully vectorized representation of agents and map elem… | 60 | 1362 | active |
| Sense-X/Co-DETR Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg… | 32 | 1357 | stable |
| ant-research/CoDeF CoDeF is the official PyTorch implementation of Content Deformation Fields, a video representation combining a canonical content field and … | 28 | 4846 | maintenance |
| mega-sam/mega-sam MegaSaM is a research codebase implementing a deep visual SLAM system that estimates camera parameters and consistent depth maps from casua… | 48 | 1355 | active |
| lxtGH/OMG-Seg Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le… | 47 | 1354 | active |
| yinguobing/head-pose-estimation A Python library for realtime human head pose estimation using ONNX Runtime and OpenCV. It combines face detection (SCRFD), 68-point facial… | 23 | 1353 | stable |
| PKU-VCL-3DV/SLAM3R SLAM3R is a real-time dense 3D scene reconstruction system that regresses 3D points from monocular RGB video using feed-forward neural netw… | 42 | 1344 | active |
| AI-FanGe/OpenAIglasses_for_Navigation An open Python framework for an AI-powered smart glasses navigation system for visually impaired users, built around an ESP32-CAM client st… | 39 | 1343 | active |
| claritylab/lucida Lucida is an open-source speech and vision based intelligent personal assistant inspired by Sirius. It orchestrates modular back-end micros… | 32 | 4782 | maintenance |