domain: computer-vision
2316 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| OpnTec/mvisc MVISC (Mobile Visual Classification) is an application that identifies and classifies individual animals from photos using computer vision,… | 32 | 1384 | experimental |
| lizhihao6/Sparc3D Sparc3D is the official implementation of a research framework for high-resolution 3D shape modeling, combining a sparse deformable marchin… | 31 | 1353 | experimental |
| DLYuanGod/TinyGPT-V TinyGPT-V is an efficient multimodal large language model built on small backbones (Phi-2 2.7B), combining vision and language capabilities… | 55 | 1316 | experimental |
| jlsutherland/doc2text doc2text is a Python library that extracts high-quality text from poorly scanned PDFs by correcting resolution, cropping, and skew before O… | 32 | 1278 | experimental |
| farzaa/gemini-bball A demo project from a viral tweet that uses Google's Gemini API to analyze basketball video frames, with an OpenCV-based visualizer. The co… | 31 | 1166 | experimental |
| geohot/twitchslam A toy monocular SLAM (Simultaneous Localization and Mapping) implementation written in Python during livestreams. It extracts features from… | 32 | 1003 | experimental |
| facebookresearch/Detectron Facebook AI Research's Python software system implementing state-of-the-art object detection algorithms such as Mask R-CNN, RetinaNet, and … | 10 | 26358 | abandoned |
| HumanSignal/labelImg LabelImg is a graphical image annotation tool written in Python with a Qt interface for drawing bounding boxes on images. It saves annotati… | 10 | 25060 | abandoned |
| apache/mxnet Apache MXNet is a deep learning framework offering a hybrid front-end that mixes imperative (Gluon) and symbolic programming, with a dynami… | 10 | 20811 | abandoned |
| jcjohnson/neural-style A Torch (Lua) implementation of the Gatys et al. neural style transfer algorithm, which combines the content of one image with the artistic… | 32 | 18284 | abandoned |
| wangshub/wechat_jump_game A Python script that plays the WeChat mini-game 'Jump Jump' (跳一跳) automatically by capturing Android screenshots via ADB, using image recog… | 23 | 13838 | abandoned |
| facebookresearch/AnimatedDrawings A Python tool from Meta AI that automatically animates children's drawings of human figures, implementing the algorithm from the paper 'A M… | 10 | 12827 | abandoned |
| react-native-camera/react-native-camera A camera component library for React Native providing photo/video capture, barcode scanning, and face detection. It is now deprecated in fa… | 10 | 9628 | abandoned |
| eduardolundgren/tracking.js tracking.js is a lightweight (~7 KB core) JavaScript library that brings computer vision algorithms like color tracking, object tracking, a… | 66 | 9465 | abandoned |
| facebookresearch/maskrcnn-benchmark A fast, modular PyTorch reference implementation of instance segmentation and object detection algorithms including Mask R-CNN, Faster R-CN… | 10 | 9360 | abandoned |
| rbgirshick/py-faster-rcnn A Python reimplementation of the Faster R-CNN object detection model built on a fork of Fast R-CNN and Caffe. The repository is officially … | 32 | 8289 | abandoned |
| jwyang/faster-rcnn.pytorch A pure PyTorch implementation of Faster R-CNN for object detection, supporting multi-image batch training and multi-GPU training with sever… | 32 | 7858 | abandoned |
| fchollet/deep-learning-models A deprecated collection of Keras code and pre-trained weights for popular deep learning image classification models such as VGG16, VGG19, R… | 23 | 7348 | abandoned |
| facebookresearch/DensePose DensePose is a research library from Facebook AI that maps all human pixels in 2D RGB images to a 3D surface-based model of the human body … | 10 | 7259 | abandoned |
| Hironsan/BossSensor A desktop application that uses a webcam and a trained CNN classifier to detect when a specific person (your boss) approaches, automaticall… | 32 | 6288 | abandoned |
| yahoo/open_nsfw A Python library from Yahoo that runs a Caffe deep neural network to classify images as Not Suitable for Work (NSFW), outputting a probabil… | 10 | 6012 | abandoned |
| oarriaga/face_classification A Python project providing real-time face detection with emotion and gender classification using a Keras CNN trained on fer2013 and IMDB da… | 32 | 5735 | abandoned |
| salesforce/BLIP PyTorch implementation of BLIP, a vision-language pre-training model for unified image understanding and text generation tasks. The reposit… | 10 | 5717 | abandoned |
| karpathy/neuraltalk2 NeuralTalk2 is a Torch (Lua) implementation of image captioning that uses a CNN (VGGNet) followed by an RNN language model to generate text… | 32 | 5592 | abandoned |
| karpathy/neuraltalk NeuralTalk is a Python+numpy implementation of Multimodal Recurrent Neural Networks that generate natural-language descriptions of images. … | 32 | 5503 | abandoned |
| dm77/barcodescanner Android library projects providing easy-to-use, extensible barcode scanner views based on ZXing and ZBar. It wraps camera preview and decod… | 10 | 5430 | abandoned |
| landing-ai/vision-agent VisionAgent is a Python library from LandingAI that takes a natural-language prompt plus an image or video and automatically selects approp… | 54 | 5296 | abandoned |
| david-gpu/srez A deep learning project that performs 4x image super-resolution on 16x16 images using a DCGAN-based architecture with ResNet generator modu… | 10 | 5270 | abandoned |
| justadudewhohacks/opencv4nodejs Node.js bindings to the native OpenCV 3 and OpenCV 4 libraries, including OpenCV-contrib modules, with both synchronous and asynchronous AP… | 23 | 5048 | abandoned |
| liuzhuang13/DenseNet Reference implementation of DenseNet (Densely Connected Convolutional Networks), the CVPR 2017 Best Paper Award-winning CNN architecture, w… | 32 | 4869 | abandoned |
| idealo/image-super-resolution A Python library providing Keras implementations of Residual Dense and Adversarial Networks for single image super-resolution, including pr… | 10 | 4818 | abandoned |
| NMAC427/SwiftOCR SwiftOCR is a fast and simple OCR library written in Swift that uses a neural network for image recognition. It is optimized for recognizin… | 23 | 4632 | abandoned |
| accord-net/framework Accord.NET is a C# framework for .NET providing machine learning, statistics, computer vision, image and audio processing, and general scie… | 10 | 4535 | abandoned |
| microsoft/VoTT VoTT (Visual Object Tagging Tool) is an open-source Electron application for annotating and labeling images and video frames to build objec… | 10 | 4428 | abandoned |
| fizyr/keras-retinanet A Keras/TensorFlow implementation of the RetinaNet object detection model with focal loss, supporting training and inference on custom data… | 23 | 4383 | abandoned |
| NVIDIA/DIGITS DIGITS is a web application for training deep learning models on GPUs, supporting frameworks like Caffe, Torch, and TensorFlow with a brows… | 10 | 4177 | abandoned |
| junyanz/iGAN iGAN is a research application implementing interactive image generation with generative adversarial networks, letting users draw strokes t… | 32 | 4005 | abandoned |
| mapillary/OpenSfM OpenSfM is a Python Structure-from-Motion library that reconstructs camera poses and 3D scenes from multiple images, including feature dete… | 67 | 3795 | abandoned |
| rmtheis/tess-two A fork of Tesseract Tools for Android providing Java APIs and build files for the Tesseract OCR and Leptonica image processing libraries on… | 10 | 3764 | abandoned |
| endernewton/tf-faster-rcnn A TensorFlow implementation of the Faster R-CNN object detection framework, supporting VGG16, ResNet, and MobileNet backbones trained and e… | 23 | 3646 | abandoned |
| rbgirshick/fast-rcnn Fast R-CNN is Ross Girshick's ICCV 2015 framework for fast object detection with deep convolutional networks, written in Python and C++/Caf… | 32 | 3460 | abandoned |
| bijection/sistine Project Sistine is a proof-of-concept Python application that turns a MacBook into a touchscreen using a $1 mirror rig in front of the buil… | 32 | 3426 | abandoned |
| oculix-org/SikuliX1 SikuliX1 is the historical Java-based visual automation tool that uses OpenCV image recognition to locate on-screen GUI elements and drive … | 78 | 3242 | abandoned |
| LeeJunHyun/Image_Segmentation A PyTorch implementation of four U-Net variants for image segmentation: U-Net, R2U-Net, Attention U-Net, and Attention R2U-Net. It includes… | 32 | 3101 | abandoned |
| facebookresearch/deepmask A Torch (Lua) implementation of the DeepMask and SharpMask object proposal algorithms from Facebook AI Research. It generates class-agnosti… | 10 | 3099 | abandoned |
| CharlesShang/FastMaskRCNN A TensorFlow implementation of Mask R-CNN for instance segmentation, reproducing the paper by Kaiming He et al. It includes ROIAlign, a COC… | 32 | 3082 | abandoned |
| haxiomic/GPU-Fluid-Experiments A cross-platform GPU-accelerated fluid simulation experiment written in Haxe, with a browser-based interactive demo. It renders real-time f… | 32 | 3066 | abandoned |
| rhsimplex/image-match A Python library for computing image signatures and searching for near-duplicate images across corpora of billions of images, backed by Ela… | 23 | 2978 | abandoned |
| xdspacelab/openvslam OpenVSLAM is a versatile visual SLAM (Simultaneous Localization and Mapping) framework supporting monocular, stereo, and RGBD cameras, incl… | 10 | 2976 | abandoned |
| ryankiros/neural-storyteller A recurrent neural network research project that generates short stories about images by combining skip-thought vectors, visual-semantic em… | 32 | 2958 | abandoned |
| jetpacapp/DeepBeliefSDK Jetpac's Deep Belief SDK, a cross-platform image recognition framework implementing the AlexNet convolutional neural network architecture, … | 32 | 2853 | abandoned |
| isl-org/ZoeDepth ZoeDepth is a PyTorch library implementing metric depth estimation from a single image, combining relative and metric depth approaches with… | 10 | 2839 | abandoned |
| ShaoqingRen/faster_rcnn A MATLAB re-implementation of Faster R-CNN, a deep-learning object detection framework combining a Region Proposal Network with a detection… | 32 | 2835 | abandoned |
| sightmachine/SimpleCV SimpleCV is an open-source Python framework that wraps OpenCV and other vision libraries behind a simple, readable API for cameras, image m… | 32 | 2730 | abandoned |
| allanzelener/YAD2K YAD2K is a 90% Keras / 10% TensorFlow implementation and converter for YOLO_v2 object detection models. It converts Darknet cfg and weight … | 32 | 2728 | abandoned |
| NVlabs/MUNIT Official PyTorch implementation of MUNIT (ECCV 2018), a GAN-based model for multimodal unsupervised image-to-image translation. The reposit… | 23 | 2707 | abandoned |
| cloud-annotations/cloud-annotations Cloud Annotations is a collaborative open-source image annotation tool for creating labeled training datasets for object detection and clas… | 23 | 2678 | abandoned |
| Guikunzhi/BeautifyFaceDemo An iOS demo app showing realtime face beautification using a custom GPUImageBeautifyFilter built on the GPUImage framework. It can be appli… | 32 | 2478 | abandoned |
| rbgirshick/rcnn The original R-CNN (Region-based Convolutional Neural Networks) object detection system from UC Berkeley, released as research code accompa… | 23 | 2416 | abandoned |
| tryolabs/luminoth Luminoth is an open-source Python toolkit for computer vision built on TensorFlow and Sonnet, focused on object detection with Faster R-CNN… | 23 | 2402 | abandoned |
| andelf/fuck12306 A Python demo project that attempts to recognize 12306 (China Railway) image CAPTCHAs using image preprocessing, pytesseract OCR, and Baidu… | 32 | 2400 | abandoned |
| kaishengtai/neuralart A Torch7 (Lua) implementation of the paper 'A Neural Algorithm of Artistic Style' for neural style transfer, applying the style of one imag… | 32 | 2397 | abandoned |
| DanielRapp/doppler A JavaScript library that detects hand motion using the doppler effect, emitting an inaudible ~20kHz tone from speakers and measuring frequ… | 32 | 2393 | abandoned |
| facebookarchive/fb.resnet.torch A Torch (Lua) implementation of ResNet residual networks for image classification, with training scripts for ImageNet and pretrained models… | 10 | 2360 | abandoned |
| microsoft/Deep3DFaceReconstruction A TensorFlow implementation of a weakly-supervised CNN method for reconstructing accurate 3D face models (shape, texture, pose, landmarks) … | 32 | 2353 | abandoned |
| colmap/glomap GLOMAP is a global structure-from-motion pipeline for image-based 3D reconstruction that takes a COLMAP database as input and outputs a COL… | 10 | 2351 | abandoned |
| isl-org/DPT DPT (Dense Prediction Transformers) is Intel's implementation of Vision Transformers for dense prediction tasks, providing pretrained model… | 10 | 2334 | abandoned |
| openai/improved-gan Reference implementation of the techniques from the OpenAI paper 'Improved Techniques for Training GANs' (arXiv:1606.03498), including semi… | 10 | 2333 | abandoned |
| card-io/card.io-iOS-SDK card.io is an iOS SDK that provides fast, easy credit card number scanning using the device camera in mobile apps. It ships as a static lib… | 10 | 2287 | abandoned |
| kongqw/OpenCVForAndroid An Android sample library built on OpenCV 3.2.0 demonstrating object detection (face, eyes, smile, body), object tracking with the CamShift… | 32 | 2228 | abandoned |
| rmtheis/android-ocr An experimental Android app that performs optical character recognition on images captured with the device camera, running the Tesseract OC… | 10 | 2226 | abandoned |
| facebookarchive/Surround360 Facebook's open-source hardware and software system for capturing and rendering stereoscopic 3D 360-degree video for VR viewing. It include… | 10 | 2185 | abandoned |
| andersbll/neural_artistic_style A Python command-line implementation of the Neural Algorithm of Artistic Style (Gatys et al., 2015) that transfers the style of one image o… | 32 | 2168 | abandoned |
| facebookresearch/SparseConvNet A PyTorch library for training Submanifold Sparse Convolutional Networks, providing spatially-sparse convolutions that operate efficiently … | 10 | 2145 | abandoned |
| neuralmagic/sparseml SparseML is a Python library for applying sparsification recipes (pruning, quantization, sparsity) to neural networks with a few lines of c… | 10 | 2145 | abandoned |
| handtracking-io/yoha Yoha is a browser-based hand tracking engine built on TensorFlow.js that detects 21 2D hand landmarks, hand presence, left/right orientatio… | 23 | 2117 | abandoned |
| zk00006/OpenTLD OpenTLD (Predator) is a MATLAB implementation of the Tracking-Learning-Detection algorithm for real-time 2D tracking of a single unknown ob… | 32 | 2101 | abandoned |
| openai/image-gpt OpenAI's official code and pre-trained models for iGPT (image GPT), a GPT-2-style transformer adapted to generate and classify images as pi… | 10 | 2096 | abandoned |
| ajbrock/Neural-Photo-Editor A GUI application for editing natural photos using generative neural networks (GANs/VAEs), implementing the Introspective Adversarial Netwo… | 32 | 2074 | abandoned |
| mitsuba-renderer/mitsuba2 Mitsuba 2 is a research-oriented, retargetable physically-based rendering system written in C++17, supporting CPU, SIMD-vectorized, and GPU… | 23 | 2072 | abandoned |
| mapbox/robosat RoboSat is an end-to-end Python pipeline for semantic segmentation and feature extraction from aerial and satellite imagery, identifying fe… | 63 | 2064 | abandoned |
| moaazsidat/react-native-qrcode-scanner A plug-and-play QR code scanner component for React Native, built on top of react-native-camera, that also works as a generic barcode scann… | 10 | 2035 | abandoned |
| card-io/card.io-Android-SDK card.io is an Android SDK that provides fast credit card scanning using the device camera, with OCR-based number recognition and a manual e… | 10 | 1995 | abandoned |
| keras-team/keras-applications Keras Applications provides reference implementations of popular deep learning models (e.g., ResNet, VGG, MobileNet) for use with Keras. Th… | 10 | 1990 | abandoned |
| facebookresearch/video-nonlocal-net A research codebase from Facebook AI implementing Non-local Neural Networks for video classification, built on the Caffe2 framework. It rep… | 10 | 1990 | abandoned |
| openai/pixel-cnn A Python/TensorFlow implementation of PixelCNN++, an autoregressive generative model for images with discretized logistic mixture likelihoo… | 10 | 1960 | abandoned |
| facebookresearch/ResNeXt A Torch (Lua) implementation of the ResNeXt architecture from the paper 'Aggregated Residual Transformations for Deep Neural Networks', bui… | 10 | 1924 | abandoned |
| justadudewhohacks/face-recognition.js A Node.js wrapper around dlib providing face detection, face recognition, and face landmark detection with JavaScript and TypeScript APIs. … | 32 | 1923 | abandoned |
| jakeret/tf_unet A generic U-Net implementation built on TensorFlow 1.x for training image segmentation models on arbitrary imaging data. Originally develop… | 23 | 1907 | abandoned |
| dlazaro66/QRCodeReaderView An Android library that wraps a modified ZXING Barcode Scanner to provide a drop-in camera view for QR code detection. It notifies listener… | 32 | 1900 | abandoned |
| erikwijmans/Pointnet2_PyTorch A PyTorch implementation of Pointnet2/Pointnet++ for deep learning on point clouds, including custom CUDA ops for GPU-only execution. It su… | 60 | 1812 | abandoned |
| ruotianluo/pytorch-faster-rcnn A PyTorch 1.0 implementation of the Faster R-CNN object detection framework, based on Xinlei Chen's tf-faster-rcnn, supporting VGG16, ResNe… | 23 | 1811 | abandoned |
| NVlabs/few-shot-vid2vid A PyTorch implementation of few-shot photorealistic video-to-video translation from NVIDIA Research (NeurIPS 2019), generating realistic vi… | 32 | 1798 | abandoned |
| traveller59/second.pytorch A PyTorch implementation of the SECOND 3D object detector for LiDAR point clouds, supporting KITTI and NuScenes datasets with sparse convol… | 32 | 1777 | abandoned |
| longcw/faster_rcnn_pytorch A PyTorch re-implementation of Faster R-CNN for object detection, based on the original py-faster-rcnn and TFFRCNN projects. It supports tr… | 32 | 1776 | abandoned |
| gram-ai/capsule-networks A barebones CUDA-enabled PyTorch implementation of the CapsNet architecture from the NIPS 2017 paper 'Dynamic Routing Between Capsules' by … | 32 | 1754 | abandoned |
| TuSimple/mx-maskrcnn An MXNet implementation of the Mask R-CNN instance segmentation model, built on top of the mx-rcnn Faster R-CNN codebase. It includes train… | 32 | 1753 | abandoned |
| dog-qiuqiu/MobileNet-Yolo A collection of ultra-lightweight YOLOv3-based object detection models (MobileNetV2-YOLOv3-Lite/Nano, YoloFace) designed for mobile and emb… | 32 | 1744 | abandoned |
| sikuli/sikuli Sikuli is a visual GUI automation tool that uses screenshot matching to find and interact with on-screen elements, originally developed as … | 32 | 1727 | abandoned |
| zhongfenglee/IDCardRecognition An Objective-C library/demo for iOS that recognizes Chinese mainland second-generation ID cards via the camera, automatically extracting na… | 32 | 1717 | abandoned |