function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ogkalu2/comic-translate An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language… | 90 | 2911 | active |
| bytedeco/javacpp-presets JavaCPP Presets provides Java bindings for commonly used native C++ libraries such as OpenCV, FFmpeg, and many others, built on the JavaCPP… | 86 | 2850 | active |
| nut-tree/nut.js nut.js is a cross-platform native UI automation and testing library for Node.js/TypeScript that controls mouse, keyboard, screen, and windo… | 23 | 2847 | active |
| QwenLM/Qwen-MM-Plugins A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.… | 57 | 2777 | active |
| apple/turicreate Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj… | 10 | 11159 | maintenance |
| AmberSahdev/Open-Interface Open Interface is a cross-platform desktop application that lets users control their computer using natural language requests processed by … | 65 | 2715 | active |
| Project-N-E-K-O/N.E.K.O Project N.E.K.O. is an open-source AI companion application — a proactive catgirl-style AI that lives on your desktop, initiates interactio… | 85 | 2681 | active |
| MrGiovanni/UNetPlusPlus Official implementation of UNet++, a nested U-Net architecture for medical image segmentation, in both Keras and PyTorch. It redesigns skip… | 77 | 2679 | stable |
| phillipi/pix2pix The original Torch (Lua) implementation of pix2pix, a conditional GAN for image-to-image translation tasks such as synthesizing photos from… | 32 | 10652 | maintenance |
| heshengtao/super-agent-party Super Agent Party is a self-hosted, all-in-one AI desktop companion that combines VRM-based virtual characters, agent skills, MCP tool supp… | 87 | 2607 | active |
| VITA-MLLM/VITA VITA is an open-source interactive omni multimodal large language model (VITA-1.5) that supports real-time vision and speech interaction, s… | 29 | 2534 | active |
| facebookresearch/pifuhd PIFuHD is a PyTorch implementation of a CVPR 2020 research model that reconstructs high-resolution 3D human body meshes from a single 2D im… | 10 | 9737 | maintenance |
| wolny/pytorch-3dunet A PyTorch implementation of 3D U-Net and its variants (residual, squeeze-and-excitation) for volumetric semantic segmentation, with 2D U-Ne… | 63 | 2416 | active |
| HuCaoFighting/Swin-Unet Official PyTorch implementation of Swin-Unet, a U-shaped pure Transformer model for medical image segmentation, published at ECCV 2022 Medi… | 41 | 2416 | stable |
| ossappscollective/OSS-DocumentScanner OSS Document Scanner is a free, open-source, privacy-focused mobile app for scanning documents with automatic edge detection, editing, OCR,… | 91 | 2385 | active |
| ailia-ai/ailia-models A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,… | 77 | 2385 | active |
| zai-org/GLM-V GLM-V is the open-source repository for Zhipu AI's GLM-4.6V, GLM-4.5V, and GLM-4.1V-Thinking vision-language models, which perform versatil… | 60 | 2370 | active |
| Turbo1123/roubao Roubao is an open-source AI phone automation assistant for Android, built natively in Kotlin and powered by vision-language models. It runs… | 53 | 2325 | active |
| PKU-YuanGroup/MoE-LLaVA MoE-LLaVA is an open-source Mixture-of-Experts based sparse large vision-language model, released with the MoE-Tuning training strategy fro… | 31 | 2322 | active |
| apple/ml-ferret Apple's Ferret, an end-to-end multimodal large language model (MLLM) that accepts any-form referring and grounds anything in its responses,… | 27 | 8674 | maintenance |
| DanOps-1/Gpt-Agreement-Payment A Python toolkit that reverse-engineers and replays the end-to-end ChatGPT Plus/Team/Pro subscription payment flow (Stripe Checkout, PayPal… | 53 | 2225 | active |
| NVlabs/MambaVision MambaVision is NVIDIA's official PyTorch implementation of a hybrid Mamba-Transformer vision backbone, published at CVPR 2025. It provides … | 49 | 2224 | active |
| ellisdg/3DUnetCNN A PyTorch library for building, training, and applying 3D U-Net convolutional neural networks for medical image segmentation. It provides c… | 45 | 2224 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2203 | active |
| kijai/ComfyUI-LivePortraitKJ ComfyUI custom nodes that integrate the LivePortrait face animation and retargeting model, supporting image-to-video, video-to-video, and n… | 23 | 2200 | active |
| cameroncooke/AXe AXe is a Swift-based CLI tool for automating and inspecting iOS Simulators on macOS using Apple's private Accessibility APIs and HID input.… | 81 | 2144 | active |
| PKU-YuanGroup/LLaVA-CoT LLaVA-CoT is an 11B vision language model with training, inference, and dataset-generation code for spontaneous step-by-step multimodal rea… | 47 | 2132 | active |
| vitoplantamura/OnnxStream A lightweight C++ inference library for ONNX models that streams weights to run large models in very little memory, accelerated by XNNPACK.… | 59 | 2086 | active |
| SizheAn/PanoHead PanoHead is the official PyTorch implementation of a CVPR 2023 paper presenting a 3D-aware GAN that synthesizes geometry-aware, view-consis… | 29 | 1956 | active |
| alibaba/EasyCV EasyCV is an all-in-one PyTorch-based computer vision toolkit from Alibaba covering self-supervised learning, vision transformers, and majo… | 32 | 1954 | active |
| microsoft/Magma Magma is Microsoft Research's foundation model for multimodal AI agents, released as an 8B vision-language model that understands images an… | 53 | 1937 | active |
| OpenTalker/video-retalking VideoReTalking is a Python research system from SIGGRAPH Asia 2022 that edits real-world talking-head videos to match a given audio track, … | 23 | 7280 | maintenance |
| IPADS-SAI/MobiAgent MobiAgent is a systematic framework for building customizable GUI agents that operate mobile phones, comprising the MobiMind agent model fa… | 60 | 1880 | active |
| NVIDIA-AI-IOT/Lidar_AI_Solution NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil… | 72 | 1867 | active |
| sarperavci/GoogleRecaptchaBypass A Python library that automatically solves Google reCAPTCHA v2 challenges in under five seconds using browser automation with DrissionPage … | 68 | 1855 | active |
| KlingAIResearch/ReCamMaster ReCamMaster is a reference implementation of a camera-controlled generative video rendering model that re-renders a single source video alo… | 44 | 1855 | active |
| NVIDIA/pix2pixHD PyTorch implementation of pix2pixHD, a conditional GAN method for synthesizing and manipulating high-resolution (2048x1024) photorealistic … | 32 | 6923 | maintenance |
| clovaai/donut Donut is an OCR-free end-to-end Transformer model for visual document understanding tasks such as document classification and information e… | 23 | 6919 | maintenance |
| GCWing/BitFun BitFun is a cross-platform desktop AI agent application with a high-performance Rust agent runtime that writes code, produces documents, an… | 82 | 1817 | active |
| yerfor/GeneFacePlusPlus GeneFace++ is the official PyTorch implementation of a NeRF-based system for generalized and stable real-time 3D talking face generation. I… | 26 | 1809 | active |
| zai-org/CogVLM CogVLM is an open-source visual language model (17B) combining a vision encoder with a pretrained language model for image understanding an… | 28 | 6744 | maintenance |
| OS-Copilot/OS-Copilot OS-Copilot is an open-source Python library for building generalist AI agents that interface with operating system elements like the web, t… | 16 | 1796 | active |
| QwenLM/Qwen-VL Official repository for Qwen-VL, Alibaba Cloud's large vision-language model family, including the pretrained Qwen-VL and instruction-tuned… | 28 | 6726 | maintenance |
| Webreaper/Damselfly Damselfly is a server-based photograph management application designed to index and search very large image collections using metadata such… | 88 | 1783 | active |
| xyTom/snippai Snippai is an AI-powered snipping tool that captures screenshots and uses AI to extract structured content such as LaTeX formulas, text, ta… | 89 | 1782 | active |
| google/automl Google Brain's AutoML repository containing implementations of AutoML models and libraries such as EfficientNet, EfficientNetV2, and Effici… | 10 | 6474 | maintenance |
| mahmoodlab/CLAM CLAM is an open-source Python toolkit for data-efficient, weakly supervised classification of whole-slide images (WSIs) in computational pa… | 39 | 1728 | active |
| facebookresearch/ConvNeXt Official PyTorch implementation of ConvNeXt, a pure convolutional neural network architecture from the CVPR 2022 paper 'A ConvNet for the 2… | 10 | 6416 | maintenance |
| AutoArk/EVA-OS EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining … | 68 | 1712 | active |
| robin-shaun/XTDrone XTDrone is a customizable UAV simulation platform built on PX4, ROS, and Gazebo, supporting multi-rotors, fixed-wing, VTOL vehicles, and ot… | 48 | 1711 | active |
| PaddlePaddle/PaddleVideo PaddleVideo is a video understanding toolkit built on PaddlePaddle, offering state-of-the-art models for action recognition, temporal actio… | 26 | 1702 | active |
| aisingapore/TagUI TagUI is a free, open-source robotic process automation (RPA) tool from AI Singapore that lets users write simple text flows to automate re… | 65 | 6325 | maintenance |
| mhamilton723/FeatUp FeatUp is a model-agnostic framework that upsamples the spatial resolution of deep neural network features by 16-32x without changing their… | 16 | 1654 | active |
| taki0112/UGATIT Official TensorFlow implementation of U-GAT-IT, an unsupervised image-to-image translation model using attention modules and adaptive layer… | 32 | 6116 | maintenance |
| ghostwright/ghost-os Ghost OS is a native macOS framework that gives AI agents full computer-use capabilities by exposing the macOS accessibility tree, Chrome D… | 61 | 1646 | active |
| CoinCheung/BiSeNet A PyTorch implementation of the BiSeNet V1 and V2 real-time semantic segmentation models, with pretrained weights for Cityscapes, COCO-Stuf… | 57 | 1638 | active |
| yoshitomo-matsubara/torchdistill torchdistill is a modular, configuration-driven PyTorch framework for knowledge distillation and general deep learning experiments, requiri… | 86 | 1629 | active |
| ZiqiaoPeng/SyncTalk SyncTalk is the official PyTorch implementation of a CVPR 2024 paper that synthesizes speech-driven, synchronized talking head videos using… | 46 | 1626 | active |
| open-gigaai/giga-world-0 GigaWorld-0 is a unified world model framework that acts as a data engine for Vision-Language-Action (VLA) learning in embodied AI. It comb… | 41 | 1612 | active |
| dmlc/gluon-cv GluonCV is a deep learning toolkit providing state-of-the-art computer vision model implementations with 170+ pre-trained models. It suppor… | 23 | 5916 | maintenance |
| pq-yang/MatAnyone MatAnyone is a CVPR 2025 human video matting framework that extracts alpha mattes of target people from video using consistent memory propa… | 54 | 1605 | active |
| ml4a/ml4a ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li… | 32 | 1602 | active |
| Roy3838/Observer Observer AI is a desktop application for building micro-agents that observe screen, camera, microphone, and audio inputs, process them with… | 87 | 1601 | active |
| semperai/amica Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. … | 32 | 1593 | active |
| Tencent-Hunyuan/HunyuanWorld-Voyager HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ… | 52 | 1590 | active |
| Drexubery/ViewCrafter ViewCrafter is a research codebase that uses video diffusion models to synthesize high-fidelity novel views of scenes from a single or spar… | 49 | 1587 | active |
| BloodAxe/pytorch-toolbelt A Python library of PyTorch extensions providing building blocks for fast R&D prototyping, including encoder-decoder architectures, special… | 44 | 1574 | active |
| byjlw/video-analyzer A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs … | 60 | 1559 | active |
| kritiksoman/GIMP-ML GIMP-ML is a set of Python plugins that bring computer vision and deep learning models into the GNU Image Manipulation Program (GIMP). It p… | 23 | 1553 | active |
| tianrun-chen/SAM-Adapter-PyTorch A PyTorch library that adapts Meta AI's Segment Anything Model (SAM, SAM2, SAM3) to underperforming downstream segmentation tasks using lig… | 67 | 1551 | active |
| cchen156/Learning-to-See-in-the-Dark TensorFlow implementation of 'Learning to See in the Dark' (CVPR 2018), a deep learning model that brightens very dark, short-exposure RAW … | 51 | 5565 | maintenance |
| WenmuZhou/PytorchOCR A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP… | 59 | 1523 | active |
| FeiYull/TensorRT-Alpha A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl… | 32 | 1460 | active |
| siddsachar/row-bot Row-Bot is a local-first desktop AI assistant and workbench that combines chat, durable memory, a personal knowledge graph, tool use, paren… | 81 | 1455 | active |
| ByteDance-Seed/m3-agent M3-Agent is a multimodal agent framework from ByteDance Seed that processes real-time visual and auditory inputs to build entity-centric lo… | 48 | 1445 | active |
| XPandora/PhysGaussian PhysGaussian is a research library that integrates Material Point Method (MPM) physics simulation with 3D Gaussian Splatting representation… | 55 | 1414 | active |
| nv-tlabs/GEN3C GEN3C is NVIDIA's research codebase for a generative video model that achieves precise camera control and temporal 3D consistency using a 3… | 59 | 1409 | active |
| sMythicalBird/ZenlessZoneZero-Auto A Python-based automation framework for the game Zenless Zone Zero that uses image classification, template matching, and OCR to perform au… | 22 | 1378 | active |
| anyrtcIO-Community/anyRTC-RTMP-OpenSource anyLive is an open-source cross-platform live streaming SDK from anyRTC built on a WebRTC-93 base, providing RTMP push (publishing) and RTM… | 30 | 4909 | maintenance |
| QwenLM/Qwen3-VL-Embedding Qwen3-VL-Embedding and Qwen3-VL-Reranker are state-of-the-art multimodal embedding and reranking models built on the Qwen3-VL foundation mo… | 56 | 1369 | active |
| ImprintLab/MedSegDiff MedSegDiff is a diffusion probabilistic model framework for segmenting and reconstructing organs and tissues from medical images, with a tr… | 50 | 1363 | active |
| a-real-ai/pywinassistant PyWinAssistant is an open-source agentic framework that acts as a Computer-Using-Agent, operating Windows 10/11 graphical user interfaces e… | 19 | 1343 | active |
| morettt/my-neuro An open-source AI desktop companion framework inspired by Neuro-sama, letting users build a customizable Live2D character with sub-second v… | 84 | 1342 | active |
| ZJU-REAL/ClawGUI ClawGUI is a unified Python framework for GUI agents covering the full lifecycle: online reinforcement learning training (ClawGUI-RL with G… | 70 | 1338 | active |
| GauravSingh9356/J.A.R.V.I.S A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,… | 48 | 1332 | active |
| ImprintLab/Medical-SAM-Adapter Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei… | 39 | 1322 | active |
| huawei-noah/Efficient-Computing A collection of efficient deep learning methods from Huawei Noah's Ark Lab, covering model compression, knowledge distillation, pruning, qu… | 32 | 1307 | active |
| robodhruv/visualnav-transformer Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie… | 19 | 1294 | active |
| jbarrow/commonforms CommonForms is a Python package and CLI that uses trained object-detection models (FFDNet-S/L) to automatically detect form fields in a PDF… | 65 | 1284 | active |
| dcharatan/pixelsplat pixelSplat is a PyTorch implementation of a feed-forward model that reconstructs 3D radiance fields parameterized by 3D Gaussian primitives… | 27 | 1274 | stable |
| Visual-Agent/DeepEyes DeepEyes is a research project that trains multimodal vision-language models to 'think with images' using end-to-end reinforcement learning… | 44 | 1271 | active |
| amaiya/ktrain ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul… | 25 | 1268 | active |
| nv-tlabs/Difix3D Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau… | 32 | 1266 | active |
| stardist/stardist StarDist is a Python library for object detection and instance segmentation in 2D and 3D microscopy images using star-convex shapes, built … | 60 | 1255 | stable |
| metadriverse/metadrive MetaDrive is an open-source, lightweight driving simulator built for AI and autonomy research, supporting compositional scene synthesis and… | 39 | 1235 | active |
| facebookresearch/home-robot HomeRobot is an open-source robotics stack from Meta AI for mobile manipulation tasks on low-cost hardware like the Hello Robot Stretch. It… | 23 | 1234 | active |
| MoonshotAI/Kimi-VL Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo… | 33 | 1224 | active |
| mrousavy/react-native-fast-tflite A high-performance TensorFlow Lite library for React Native built on Nitro Modules, using the low-level C/C++ TFLite core API with zero-cop… | 84 | 1222 | active |
| frotms/PaddleOCR2Pytorch A PyTorch port of PaddleOCR that lets you run PaddleOCR-trained models (detection, recognition, and document structure parsing) without the… | 73 | 1205 | active |
| deepseek-ai/DeepSeek-VL DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf… | 25 | 4175 | maintenance |