function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| leftthomas/SRGAN A PyTorch implementation of SRGAN, the CVPR 2017 generative adversarial network for photo-realistic single-image super-resolution. It inclu… | 32 | 1248 | maintenance |
| lfz/DSB2017 The winning solution of team 'grt123' for the 2017 Data Science Bowl (DSB2017), a deep learning pipeline for detecting lung cancer from CT … | 32 | 1241 | maintenance |
| BR-IDL/PaddleViT PaddleViT is a collection of state-of-the-art Vision Transformer and MLP model implementations for PaddlePaddle 2.1+, covering image classi… | 23 | 1238 | maintenance |
| openseg-group/openseg.pytorch Official PyTorch implementations of semantic segmentation models OCNet, OCRNet, and SegFix, achieving state-of-the-art results on benchmark… | 23 | 1237 | maintenance |
| jfzhang95/pytorch-video-recognition A PyTorch library implementing C3D, R3D, and R2Plus1D models for video action recognition, with training scripts for UCF101 and HMDB51 data… | 32 | 1236 | maintenance |
| ml4a/ml4a-ofx A collection of openFrameworks applications in C++ for real-time interactive machine learning, aimed at artists and creative coders. It inc… | 23 | 1223 | maintenance |
| facebookresearch/House3D House3D is a virtual 3D environment of over 45k fully annotated indoor scenes from the SUNCG dataset, built for training embodied AI agents… | 10 | 1200 | maintenance |
| VITA-Group/DeblurGANv2 Official PyTorch implementation of DeblurGAN-v2, an ICCV 2019 relativistic conditional GAN for single-image motion deblurring with a Featur… | 32 | 1192 | maintenance |
| aitorzip/DeepGTAV DeepGTAV is a C++ plugin for Grand Theft Auto V that converts the game into a vision-based self-driving car research environment. It expose… | 32 | 1186 | maintenance |
| cpvrlab/ImagePlay ImagePlay is an open-source desktop application for rapid prototyping of image processing algorithms, combining over 70 individual image pr… | 23 | 1173 | maintenance |
| Jinnrry/RobotHelper RobotHelper is an Android automation script framework written in Java, providing common building blocks like screen capture, image-based po… | 23 | 1136 | maintenance |
| yu-takagi/StableDiffusionReconstruction Research codebase reproducing Takagi and Nishimoto's CVPR 2023 method for reconstructing images a person viewed from fMRI brain activity us… | 30 | 1127 | maintenance |
| snap-research/EfficientFormer A PyTorch implementation of EfficientFormer and EfficientFormerV2, efficient vision transformer model families designed to run at MobileNet… | 32 | 1116 | maintenance |
| Res2Net/Res2Net-PretrainedModels Official PyTorch implementation of Res2Net, a multi-scale CNN backbone architecture published in TPAMI, with ImageNet-pretrained model weig… | 32 | 1114 | maintenance |
| andyzeng/visual-pushing-grasping PyTorch reference implementation of Visual Pushing and Grasping (VPG), which trains robotic agents via self-supervised deep reinforcement l… | 32 | 1109 | maintenance |
| qiucheng025/zao- A Python deep learning tool that identifies and swaps faces in images and videos, with extract, train, and convert workflows plus an option… | 32 | 1104 | maintenance |
| PaddlePaddle/Paddle.js Paddle.js is a browser-based deep learning inference engine for Baidu PaddlePaddle models, running via WebGL, WebGPU, or WebAssembly backen… | 23 | 1103 | maintenance |
| fyu/drn A PyTorch library implementing Dilated Residual Networks (DRN), which combine dilated convolutions with residual networks for image classif… | 32 | 1102 | maintenance |
| jacobgil/vit-explain A PyTorch library implementing explainability methods for Vision Transformers, including Attention Rollout and Gradient Attention Rollout. … | 32 | 1098 | maintenance |
| maelfabien/Multimodal-Emotion-Recognition A real-time multimodal emotion recognition web app built with Flask that analyzes emotions from text, audio, and video inputs using deep le… | 32 | 1089 | maintenance |
| ethanhe42/channel-pruning Reference implementation of the ICCV 2017 channel pruning method for accelerating very deep convolutional neural networks, using LASSO regr… | 23 | 1088 | maintenance |
| deepmedic/deepmedic DeepMedic is an efficient multi-scale 3D convolutional neural network for segmenting 3D medical scans such as MRI and CT. It is a Python-ba… | 23 | 1063 | maintenance |
| HRNet/HRNet-Image-Classification Official PyTorch implementation and training code for HRNet (High-Resolution Network) image classification models on ImageNet. It provides … | 23 | 1056 | maintenance |
| piergiaj/pytorch-i3d A PyTorch port of DeepMind's I3D (Inflated 3D ConvNet) models pretrained on the Kinetics dataset for video action recognition. It includes … | 32 | 1054 | maintenance |
| 4uiiurz1/pytorch-nested-unet A PyTorch implementation of the UNet++ (Nested U-Net) architecture for image segmentation, based on the paper 'UNet++: A Nested U-Net Archi… | 32 | 1045 | maintenance |
| goberoi/faceit A Python script that simplifies swapping faces in videos using the deepfakes/faceswap library, with training data sourced from YouTube vide… | 32 | 1032 | maintenance |
| qubvel/ttach TTAch is a Python library for image test time augmentation (TTA) with PyTorch. It wraps existing models to apply augmentations like flips, … | 23 | 1030 | maintenance |
| Gumpest/YOLOv5-Multibackbone-Compression A YOLOv5-based toolbox for swapping in lightweight or high-accuracy backbones (TPH-YOLOv5, GhostNet, ShuffleNetV2, MobileNetV3-Small, Effic… | 32 | 1020 | maintenance |
| snap-research/NeROIC Official PyTorch implementation of NeROIC, a neural method for capturing 3D object geometry and material from online image collections and … | 32 | 1012 | maintenance |
| xiaoyufenfei/Efficient-Segmentation-Networks A PyTorch reference implementation collection of lightweight, real-time semantic segmentation models such as ENet, ERFNet, LEDNet, Fast-SCN… | 32 | 1008 | maintenance |
| dimensionalOS/dimos DimOS is a Python SDK and agent-native operating system for generalist robotics, letting users command humanoids, quadrupeds, and drones in… | 86 | 4431 | experimental |
| cbh123/narrator A Python app that watches your webcam and generates David Attenborough-style narration of what it sees, using GPT vision models and ElevenL… | 69 | 4426 | experimental |
| IliasHad/edit-mind Edit Mind is a local-first video knowledge base that indexes video libraries with multi-modal AI analysis (Whisper transcription, YOLO obje… | 76 | 1789 | experimental |
| jlsutherland/doc2text doc2text is a Python library that extracts high-quality text from poorly scanned PDFs by correcting resolution, cropping, and skew before O… | 32 | 1278 | experimental |
| ddupont808/GPT-4V-Act GPT-4V-Act is a multimodal AI agent that combines GPT-4V(ision) with a Chromium browser, using Set-of-Mark prompting and a DOM auto-labeler… | 27 | 1058 | experimental |
| fchollet/deep-learning-models A deprecated collection of Keras code and pre-trained weights for popular deep learning image classification models such as VGG16, VGG19, R… | 23 | 7348 | abandoned |
| karpathy/neuraltalk NeuralTalk is a Python+numpy implementation of Multimodal Recurrent Neural Networks that generate natural-language descriptions of images. … | 32 | 5503 | abandoned |
| liuzhuang13/DenseNet Reference implementation of DenseNet (Densely Connected Convolutional Networks), the CVPR 2017 Best Paper Award-winning CNN architecture, w… | 32 | 4869 | abandoned |
| Greenwolf/social_mapper Social Mapper is a Python 3 OSINT tool that enumerates and correlates social media profiles across sites like LinkedIn, Facebook, Twitter, … | 32 | 4073 | abandoned |
| udacity/self-driving-car-sim A Unity-based self-driving car simulator built for Udacity's Self-Driving Car Nanodegree, used to train cars to navigate road courses with … | 10 | 3985 | abandoned |
| pavelgonchar/colornet Colornet is a Python/TensorFlow neural network implementation that colorizes grayscale images, based on VGG16 features and YUV color-channe… | 32 | 3552 | abandoned |
| LeeJunHyun/Image_Segmentation A PyTorch implementation of four U-Net variants for image segmentation: U-Net, R2U-Net, Attention U-Net, and Attention R2U-Net. It includes… | 32 | 3101 | abandoned |
| ryanjay0/miles-deep Miles Deep is a C++ deep learning application built on Caffe that classifies each second of a pornographic video into six sexual act catego… | 23 | 2658 | abandoned |
| jakeret/tf_unet A generic U-Net implementation built on TensorFlow 1.x for training image segmentation models on arbitrary imaging data. Originally develop… | 23 | 1907 | abandoned |
| sikuli/sikuli Sikuli is a visual GUI automation tool that uses screenshot matching to find and interact with on-screen elements, originally developed as … | 32 | 1727 | abandoned |
| ramprs/grad-cam Official Torch (Lua) implementation of Grad-CAM, the ICCV 2017 gradient-weighted class activation mapping technique for producing visual ex… | 32 | 1668 | abandoned |
| Zehaos/MobileNet A TensorFlow implementation of Google's MobileNets, efficient convolutional neural networks for mobile vision applications, including Image… | 32 | 1659 | abandoned |
| bonlime/keras-deeplab-v3-plus A Keras implementation of the DeepLab v3+ semantic image segmentation model with pretrained weights imported from the original TensorFlow c… | 23 | 1375 | abandoned |
| MIC-DKFZ/medicaldetectiontoolkit A PyTorch framework providing 2D and 3D implementations of object detectors like Mask R-CNN, Retina Net, and Retina U-Net, tailored for med… | 32 | 1357 | abandoned |
| piiswrong/deep3d Deep3D is a CNN-based research project that automatically converts 2D images and videos into 3D by estimating per-pixel depth maps and gene… | 32 | 1299 | abandoned |
| shekkizh/FCN.tensorflow A TensorFlow implementation of Fully Convolutional Networks (FCN) for semantic segmentation, based on the reference code from the original … | 32 | 1248 | abandoned |
| Teaonly/android-eye An Android app that turns an old phone into a surveillance security camera, streaming H.264 video and G.726 audio over a built-in web serve… | 10 | 1085 | abandoned |
| alexgkendall/caffe-segnet A modified version of the Caffe deep learning framework implementing SegNet, a deep convolutional encoder-decoder architecture for semantic… | 32 | 1083 | abandoned |
| DeepLearningKit/DeepLearningKit DeepLearningKit is an open-source deep learning framework for Apple's iOS, OS X and tvOS, written in Swift and using Metal for GPU-accelera… | 32 | 1059 | abandoned |
| KevinGong2013/ChineseIDCardOCR A deprecated Swift library for optical character recognition of Chinese second-generation ID cards on iOS, using Vision and CoreML. It has … | 32 | 1025 | abandoned |
← prev page 16 / 16