function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| facebookresearch/multipathnet A Torch-7 implementation of the MultiPath Network for object detection from the BMVC 2016 paper by Facebook AI Research, also supporting Fa… | 10 | 1331 | abandoned |
| trailbehind/DeepOSM DeepOSM is a Python application that trains neural networks with TensorFlow to classify roads and features in satellite imagery using OpenS… | 32 | 1330 | abandoned |
| YonghaoHe/LFFD-A-Light-and-Fast-Face-Detector-for-Edge-Devices LFFD is a light and fast single-class object detection framework designed for edge devices, with pretrained models for face, head, pedestri… | 32 | 1323 | abandoned |
| datitran/object_detector_app A Python application that performs real-time object recognition from a webcam or HLS video stream using TensorFlow's Object Detection API a… | 32 | 1305 | abandoned |
| tianzhi0549/CTPN Reference implementation of CTPN (Connectionist Text Proposal Network) for detecting text lines in natural images, from the ECCV 2016 paper… | 32 | 1287 | abandoned |
| Prinsphield/Wechat_AutoJump A Python bot that automatically plays the WeChat 'Jump Jump' mini-game using computer vision and a CNN coarse-to-fine model to locate the p… | 32 | 1280 | abandoned |
| MicrosoftEdge/magic-mirror-demo A smart mirror IoT demo project by Microsoft Edge that displays information on a two-way mirror and recognizes registered users via facial … | 10 | 1225 | abandoned |
| CarlosGS/Cyclone-PCB-Factory Cyclone PCB Factory is a parametric, 3D-printable CNC mill design (RepRap-style) intended for milling printed circuit boards. It provides C… | 10 | 1187 | abandoned |
| Jai-wei/YOLOv8-PySide6-GUI YoloSide is a desktop GUI application built with PySide6 for running YOLOv8 object detection models. Users can load trained .pt model files… | 30 | 1171 | abandoned |
| linkedlist771/SoraWatermarkCleaner A deep learning tool that detects and removes the Sora2 watermark from AI-generated videos using a YOLO-based detector plus a restoration m… | 10 | 1149 | abandoned |
| faceair/youjumpijump A Go-based cheat bot for the WeChat 'Jump Jump' (跳一跳) mini-game that screenshots the screen, computes jump distance via image analysis, and… | 10 | 1146 | abandoned |
| eldar/pose-tensorflow A TensorFlow implementation of the DeeperCut and ArtTrack algorithms for human body pose estimation, supporting both single-person and mult… | 32 | 1141 | abandoned |
| szad670401/end-to-end-for-chinese-plate-recognition An end-to-end Chinese license plate recognition model based on MXnet, using multi-label classification. It was trained on ~500k synthetic r… | 32 | 1117 | abandoned |
| cgtinker/BlendArMocap A Blender add-on that performs markerless motion capture using Google's Mediapipe, detecting pose, hand, and face features from webcam stre… | 48 | 1078 | abandoned |
| mikebuss/MTBBarcodeScanner A lightweight Objective-C barcode scanning library for iOS built on AVFoundation, supporting single and multiple barcode detection, torch c… | 10 | 1075 | abandoned |
| tomthecarrot/arcore-for-all A modified build of Google's ARCore developer preview library that removes the official device whitelist check, enabling ARCore to run on u… | 10 | 1054 | abandoned |
| burningcl/wechat_jump_hack A Java-based bot that automatically plays WeChat's 'Jump Jump' (跳一跳) mini-game by capturing screenshots via ADB, recognizing player and tar… | 32 | 1053 | abandoned |
| facebookresearch/VMZ VMZ is a model zoo from Facebook AI's Computer Vision team providing Caffe2 and PyTorch implementations of video classification models such… | 10 | 1052 | abandoned |
| NVIDIA-AI-IOT/redtail NVIDIA Redtail provides deep learning and computer vision components for autonomous visual navigation of drones and ground vehicles, center… | 23 | 1047 | abandoned |
| digital-standard/ThreeDPoseTracker A Unity-based Windows application that estimates 3D human pose from video or webcam input using an ONNX neural network model via Unity Barr… | 23 | 1042 | abandoned |
| PRBonn/lidar-bonnetal A deep learning framework for training and deploying semantic segmentation of LiDAR point clouds using range-image representations, develop… | 10 | 1037 | abandoned |
| doug/depthjs DepthJS is a browser extension and native plugin (primarily for Chrome) that lets any web page interact with the Microsoft Kinect via JavaS… | 32 | 1001 | abandoned |
| tensorflow/models The TensorFlow Model Garden is a repository of official and community implementations of state-of-the-art machine learning models built wit… | 85 | 77652 | active |
| deepfakes/faceswap Faceswap is a free, open-source, multi-platform deepfakes tool that uses deep learning to recognize and swap faces in pictures and videos. … | 85 | 57500 | active |
| mudler/LocalAI LocalAI is an open-source, self-hosted AI inference runtime that runs LLMs, vision, speech, image, and video models behind OpenAI/Anthropic… | 93 | 48696 | active |
| huggingface/pytorch-image-models PyTorch Image Models (timm) is a Python library offering the largest collection of PyTorch image encoder/backbone architectures with 700+ p… | 93 | 37099 | active |
| 78/xiaozhi-esp32 XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen… | 88 | 29186 | active |
| ente/ente Ente is an open-source, end-to-end encrypted cloud platform with client apps for photos (a Google Photos alternative), document/credential … | 94 | 28513 | active |
| shap/shap SHAP (SHapley Additive exPlanations) is a Python library that explains the output of any machine learning model using Shapley values from g… | 90 | 25704 | active |
| junyanz/pytorch-CycleGAN-and-pix2pix Official PyTorch implementations of CycleGAN and pix2pix for paired and unpaired image-to-image translation. It includes training and testi… | 48 | 25232 | stable |
| deepseek-ai/DeepSeek-OCR DeepSeek-OCR is an open vision-language model from DeepSeek AI that researches 'contexts optical compression' - encoding long text contexts… | 45 | 23855 | active |
| Skyvern-AI/skyvern Skyvern is an open-source AI browser automation framework that uses LLMs and computer vision to interact with websites, offering a Playwrig… | 88 | 22852 | active |
| microsoft/unilm Microsoft's collection of large-scale self-supervised pre-trained models spanning tasks, 100+ languages, and modalities (text, image, layou… | 67 | 22194 | active |
| huggingface/datasets Hugging Face Datasets is a Python library providing one-line access to hundreds of thousands of public datasets on the Hugging Face Hub acr… | 98 | 21870 | stable |
| screenpipe/screenpipe Screenpipe is a source-available desktop application that continuously records your screen and audio locally, extracting text via OCR/acces… | 86 | 21244 | active |
| huggingface/candle Candle is a minimalist machine learning framework for Rust focused on performance and ease of use, with CPU and CUDA GPU support. It ships … | 73 | 20955 | active |
| QwenLM/Qwen3-VL Qwen3-VL is a series of open-weight multimodal vision-language models from Alibaba's Qwen team, available in Dense and MoE architectures wi… | 52 | 19847 | active |
| huggingface/transformers.js Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) … | 91 | 16270 | active |
| microsoft/Swin-Transformer Official PyTorch implementation of the Swin Transformer, a hierarchical vision transformer using shifted windows that serves as a general-p… | 32 | 16051 | stable |
| alibaba/MNN MNN is a lightweight, high-performance deep learning inference engine developed by Alibaba, supporting LLMs, vision, audio, and multimodal … | 93 | 15973 | active |
| duixcom/Duix-Avatar Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea… | 57 | 14871 | active |
| dlib Dlib is a modern C++ toolkit containing machine learning algorithms, deep learning tools, computer vision, linear algebra, and general-purp… | 86 | 14431 | stable |
| jacobgil/pytorch-grad-cam A PyTorch library providing state-of-the-art pixel attribution (saliency) methods like GradCAM, ScoreCAM, and AblationCAM for explainable A… | 76 | 12958 | active |
| octalmage/robotjs RobotJS is a Node.js desktop automation library for controlling the mouse and keyboard and reading the screen, with native prebuilt binarie… | 93 | 12770 | active |
| ludwig-ai/ludwig Ludwig is a declarative, low-code deep learning framework for training, fine-tuning, and deploying AI models — from LLMs to tabular, image,… | 99 | 11745 | active |
| qubvel-org/segmentation_models.pytorch A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar… | 70 | 11706 | stable |
| milesial/Pytorch-UNet A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva… | 23 | 11613 | active |
| facebookresearch/dinov3 Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t… | 59 | 11249 | active |
| voxel51/fiftyone FiftyOne is an open-source Python library and GUI app for building high-quality computer vision datasets and models. It enables visualizing… | 99 | 11042 | active |
| OpenVINO OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su… | 95 | 10740 | stable |
| autogluon/autogluon AutoGluon is an AutoML library that automates machine learning on tabular data, time series, text, and images with just a few lines of Pyth… | 90 | 10617 | active |
| thumbor/thumbor Thumbor is an open-source, on-demand image thumbnailing service written in Python. It crops, resizes, flips, and applies filters to images … | 90 | 10514 | active |
| openframeworks/openFrameworks openFrameworks is an open-source C++ toolkit for creative coding that wraps common libraries like OpenGL, OpenCV, and audio/video libraries… | 87 | 10419 | stable |
| OpenGVLab/InternVL InternVL is a family of open-source multimodal large language models (vision-language models) that combine vision transformers with LLMs to… | 37 | 10146 | active |
| open-mmlab/mmsegmentation MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat… | 23 | 9930 | stable |
| bytedance/Dolphin Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-… | 52 | 9049 | active |
| apple/ml-sharp SHARP is a Python tool from Apple that synthesizes a photorealistic 3D Gaussian splat representation from a single photograph in under a se… | 42 | 8843 | active |
| firerpa/lamda FIRERPA (lamda) is an all-in-one Android device control platform whose server runs directly on the device (root or non-root) and exposes 16… | 98 | 8243 | active |
| GetStream/Vision-Agents An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O… | 84 | 8100 | active |
| facebookresearch/SlowFast PySlowFast is a PyTorch-based open-source video understanding codebase from Facebook AI Research (FAIR). It provides implementations of sta… | 65 | 7410 | active |
| farzaa/clicky Clicky is an open-source macOS app that acts as an AI teacher living next to your cursor — it can see your screen, talk with you via voice,… | 50 | 7394 | active |
| facebookresearch/sam-3d-objects SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code… | 55 | 7322 | active |
| turanszkij/WickedEngine Wicked Engine is an open-source C++ 3D game engine with modern graphics features like ray tracing, global illumination, and physically base… | 99 | 7203 | active |
| PaddlePaddle/models PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda… | 23 | 6932 | active |
| ml5js/ml5-library ml5.js is a friendly, beginner-oriented JavaScript machine learning library for the browser, built on top of TensorFlow.js. It provides acc… | 23 | 6587 | active |
| iperov/DeepFaceLive DeepFaceLive is a real-time face-swap application for PC streaming and video calls, using trained face models (DFM) applied to webcam or vi… | 10 | 31011 | maintenance |
| PaddlePaddle/PaddleX PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti… | 92 | 6251 | active |
| om-ai-lab/VLM-R1 VLM-R1 is a framework for training R1-style large vision-language models using reinforcement learning (GRPO) on top of Qwen2.5-VL. It provi… | 63 | 6015 | active |
| PaddlePaddle/PaddleClas PaddleClas is a Python library and toolkit for image classification, recognition, and retrieval built on the PaddlePaddle deep learning fra… | 66 | 5838 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| dnhkng/GLaDOS A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,… | 64 | 5689 | active |
| microsoft/SynapseML SynapseML (formerly MMLSpark) is an open-source machine learning library built on Apache Spark that provides simple, composable, distribute… | 88 | 5240 | active |
| katanaml/sparrow Sparrow is an open-source framework for structured data extraction from documents (PDFs, images) using ML, LLMs, and Vision LLMs, with sche… | 95 | 5202 | active |
| ai-dawang/PlugNPlay-Modules A curated collection of plug-and-play deep learning modules (convolutions, attention mechanisms, downsampling, and feature fusion blocks) i… | 38 | 5105 | active |
| dmMaze/BallonsTranslator A desktop GUI application that uses deep learning to automatically translate comics and manga, combining text detection, OCR, inpainting, a… | 99 | 5065 | active |
| KaiyangZhou/deep-person-reid Torchreid is a PyTorch library for deep-learning person re-identification, supporting both image and video reid with end-to-end training an… | 50 | 4900 | stable |
| aloshdenny/reverse-SynthID A research tool that reverse-engineers Google's SynthID watermark embedded in Gemini-generated images using spectral analysis and signal pr… | 66 | 4815 | active |
| open-mmlab/mmocr MMOCR is OpenMMLab's PyTorch-based toolbox for text detection, recognition, and key information extraction. It provides a model zoo of OCR … | 23 | 4752 | active |
| Tencent/TNN TNN is a high-performance, lightweight deep learning inference framework developed by Tencent Youtu Lab, supporting mobile, desktop, and se… | 32 | 4648 | active |
| facebookresearch/vjepa2 Official PyTorch codebase and pretrained models for V-JEPA 2, a self-supervised video encoder trained on internet-scale video, plus V-JEPA … | 52 | 4527 | active |
| layumi/Person_reID_baseline_pytorch A small, friendly PyTorch baseline implementation for person and vehicle re-identification (ReID). It reproduces strong top-conference resu… | 65 | 4446 | stable |
| xlite-dev/lite.ai.toolkit A lightweight C++ toolkit providing unified APIs for 100+ pre-trained AI models across inference backends like ONNX Runtime, MNN, TensorRT,… | 74 | 4427 | active |
| open-compass/VLMEvalKit VLMEvalKit is an open-source Python toolkit for evaluating large vision-language models (LMMs/LVLMs) across 80+ benchmarks with support for… | 62 | 4359 | active |
| httprunner/httprunner HttpRunner (hrp) is an open-source, Go-based all-in-one testing framework for API testing (HTTP/HTTP2/WebSocket/RPC), load testing, and mul… | 48 | 4295 | active |
| QwenLM/Qwen2.5-Omni Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre… | 31 | 4074 | active |
| thuml/Transfer-Learning-Library TLlib is a PyTorch-based open-source library for transfer learning, covering domain adaptation, task adaptation (finetuning), and domain ge… | 23 | 3931 | active |
| NVlabs/VILA VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d… | 57 | 3857 | active |
| open-mmlab/mmpretrain MMPretrain is OpenMMLab's PyTorch-based toolbox and benchmark for image classification model pre-training, covering supervised, self-superv… | 23 | 3850 | active |
| shitagaki-lab/see-through A research framework from a SIGGRAPH 2026 paper that decomposes a single anime character illustration into up to 23 fully inpainted, semant… | 58 | 3641 | active |
| NVlabs/Eagle Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L… | 64 | 3462 | active |
| Kedreamix/Linly-Talker Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, … | 48 | 3436 | active |
| opengeos/geoai GeoAI is a Python package that integrates artificial intelligence with geospatial data analysis, built on PyTorch, Transformers, and segmen… | 89 | 3327 | active |
| deepdoctection/deepdoctection deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c… | 98 | 3248 | active |
| Beckschen/TransUNet Official PyTorch implementation of TransUNet, a U-Net-style architecture that uses a Vision Transformer encoder for medical image segmentat… | 63 | 3234 | stable |
| facebookresearch/dinov2 PyTorch implementation and pretrained models for DINOv2, a self-supervised vision transformer method from Meta AI that learns robust visual… | 68 | 13266 | maintenance |
| junyanz/CycleGAN A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I… | 32 | 12870 | maintenance |
| SharpAI/DeepCamera DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r… | 86 | 3019 | active |
| osmr/imgclsmob A research sandbox providing (re)implementations of numerous deep learning computer vision models for classification, segmentation, detecti… | 23 | 3016 | active |
| Tamsiree/RxTool RxTool is a large collection of utility classes and UI components for Android development, packaged as multiple Gradle modules (RxKit, RxUI… | 23 | 12288 | maintenance |
| sherlockchou86/VideoPipe VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates … | 54 | 2931 | active |