function: image-processing
4273 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| yu4u/age-gender-estimation A Keras/TensorFlow implementation of a convolutional neural network that estimates age and gender from face images, trained on the IMDB-WIK… | 23 | 1520 | maintenance |
| marsauto/europilot Europilot is a Python toolkit that bridges Euro Truck Simulator 2 with deep-learning frameworks, enabling end-to-end self-driving research … | 32 | 1511 | maintenance |
| tinyvision/SOLIDER SOLIDER is a semantic-controllable self-supervised learning framework that learns general human representations from massive unlabeled huma… | 32 | 1503 | maintenance |
| AITTSMD/MTCNN-Tensorflow A TensorFlow reproduction of MTCNN (Multi-task Cascaded Convolutional Networks) for joint face detection and facial landmark alignment. It … | 32 | 1502 | maintenance |
| NVIDIAGameWorks/kaolin-wisp NVIDIA Kaolin Wisp is a PyTorch library and engine for neural fields research, built on NVIDIA Kaolin Core. It provides differentiable rend… | 23 | 1496 | maintenance |
| speedinghzl/CCNet Official PyTorch implementation of CCNet, a Criss-Cross Attention network for semantic segmentation published at ICCV 2019 and TPAMI 2020. … | 32 | 1485 | maintenance |
| chengdazhi/Deformable-Convolution-V2-PyTorch A PyTorch implementation of Deformable Convolution V2 (DCNv2) custom CUDA operators, ported from the original MXNet implementation. It prov… | 32 | 1484 | maintenance |
| ruotianluo/ImageCaptioning.pytorch A PyTorch research codebase for image captioning, supporting self-critical sequence training, bottom-up features, transformer captioning mo… | 32 | 1476 | maintenance |
| faustomorales/keras-ocr A Python library packaging the CRAFT text detector and a Keras CRNN text recognition model into a high-level OCR pipeline. It supports pret… | 42 | 1473 | maintenance |
| HRNet/HigherHRNet-Human-Pose-Estimation Official PyTorch implementation of HigherHRNet, a CVPR 2020 bottom-up multi-person human pose estimation model using scale-aware high-resol… | 32 | 1464 | maintenance |
| yangyanli/PointCNN PointCNN is a deep learning framework for feature learning from 3D point clouds, applying convolution on X-transformed points to handle the… | 55 | 1433 | maintenance |
| qfgaohao/pytorch-ssd A PyTorch implementation of the SSD (Single Shot MultiBox Detector) object detection algorithm with MobileNetV1, MobileNetV2, and VGG backb… | 32 | 1429 | maintenance |
| theAIGuysCode/yolov4-deepsort A Python implementation of multi-object tracking that combines YOLOv4 object detection with the Deep SORT tracking algorithm using TensorFl… | 32 | 1428 | maintenance |
| jonnnnyw/php-phantomjs A PHP library that wraps the PhantomJS headless browser, letting PHP applications load web pages with full JavaScript support and inspect t… | 23 | 1426 | maintenance |
| open-mmlab/mmhuman3d MMHuman3D is an open-source PyTorch-based toolbox and benchmark for 3D human parametric models (e.g., SMPL, SMPL-X) in computer vision and … | 23 | 1425 | maintenance |
| rmokady/CLIP_prefix_caption Official implementation of ClipCap, a CLIP-based image captioning model that maps CLIP image encodings to a GPT-2 prefix to generate captio… | 32 | 1423 | maintenance |
| ProbableTrain/MapGenerator A browser-based tool that procedurally generates American-style city maps with images exportable as PNG or SVG and 3D models exportable as … | 32 | 1418 | maintenance |
| uzh-rpg/flightmare Flightmare is a modular quadrotor simulator from the UZH Robotics and Perception Group, composed of a decoupled Unity-based rendering engin… | 23 | 1410 | maintenance |
| chrischoy/3D-R2N2 3D-R2N2 is a PyTorch-based implementation of a recurrent neural network that reconstructs voxelized 3D models of objects from one or multip… | 32 | 1409 | maintenance |
| xuan32546/IOS13-SimulateTouch A system-wide touch event simulation library and tweak for jailbroken iOS devices running iOS 11.0-14, distributed as the ZXTouch deb packa… | 23 | 1407 | maintenance |
| PSPNet Reference implementation of the Pyramid Scene Parsing Network (PSPNet), a CVPR 2017 semantic segmentation model that won the ImageNet Scene… | 32 | 1378 | maintenance |
| tleyden/open-ocr OpenOCR is a self-hosted OCR-as-a-Service REST API built in Go on top of Tesseract, containerized with Docker. It uses RabbitMQ for scalabl… | 32 | 1373 | maintenance |
| nv-tlabs/lift-splat-shoot PyTorch implementation of Lift-Splat-Shoot (ECCV 2020), an end-to-end model that converts images from arbitrary multi-camera rigs into a bi… | 32 | 1369 | maintenance |
| wzzheng/TPVFormer TPVFormer is a CVPR 2023 research implementation of a tri-perspective view transformer for vision-based 3D semantic occupancy prediction, s… | 32 | 1364 | maintenance |
| MoyGcc/vid2avatar Vid2Avatar is the official PyTorch implementation of a CVPR 2023 method that reconstructs detailed 3D human avatars from monocular in-the-w… | 57 | 1360 | maintenance |
| mayuelala/FollowYourPose Official PyTorch implementation of Follow-Your-Pose (AAAI 2024), a pose-guided text-to-video generation model that tunes a text-to-image mo… | 21 | 1358 | maintenance |
| marcoarment/BugshotKit An iOS library that lets developers and beta testers trigger in-app bug reports via a gesture, capturing an annotated screenshot and the li… | 32 | 1352 | maintenance |
| PeizeSun/SparseR-CNN Sparse R-CNN is a PyTorch implementation (built on Detectron2) of the CVPR 2021 / PAMI 2023 paper 'End-to-End Object Detection with Learnab… | 23 | 1343 | maintenance |
| torrvision/crfasrnn Reference implementation of CRF-RNN, an ICCV 2015 semantic image segmentation method that integrates conditional random fields into a neura… | 23 | 1335 | maintenance |
| maudzung/Complex-YOLOv4-Pytorch A PyTorch implementation of Complex-YOLO, a YOLOv4-based model for real-time 3D object detection on LiDAR point clouds. It supports distrib… | 32 | 1327 | maintenance |
| jingle1267/android-utils A Java utility class library for Android development that bundles dozens of commonly used helper classes covering files, bitmaps, networkin… | 32 | 1322 | maintenance |
| vercel/modelfusion ModelFusion is a TypeScript library that provides a unified, vendor-neutral abstraction layer for integrating AI models into JavaScript and… | 10 | 1319 | maintenance |
| LinXueyuanStdio/LaTeX_OCR_PRO A deep-learning application that converts images of math formulas (printed, handwritten, and Chinese-mixed) into LaTeX code using a Seq2Seq… | 32 | 1312 | maintenance |
| senlinuc/caffe_ocr An experimental research project built on Caffe implementing CNN+BLSTM+CTC text recognition architectures, with modifications for LSTM, war… | 32 | 1305 | maintenance |
| NVlabs/DG-Net DG-Net is a PyTorch implementation of the CVPR 2019 (Oral) paper 'Joint Discriminative and Generative Learning for Person Re-identification… | 32 | 1298 | maintenance |
| apple/ml-neuman Official reference implementation of NeuMan (ECCV 2022), which reconstructs an animatable human and the background scene from a single vide… | 32 | 1287 | maintenance |
| ifzc/Shkjem A corporate website frontend for Shanghai Kejian Engineering Management, built with Vue, Element UI, and vue-cli 3. It was originally a thr… | 34 | 1281 | maintenance |
| SullyChen/Autopilot-TensorFlow A TensorFlow implementation of Nvidia's end-to-end self-driving steering angle prediction paper (arXiv 1604.07316) with some modifications.… | 32 | 1276 | maintenance |
| j96w/DenseFusion DenseFusion is the official PyTorch implementation of the paper '6D Object Pose Estimation by Iterative Dense Fusion', which estimates the … | 32 | 1276 | maintenance |
| TRI-ML/packnet-sfm Official PyTorch implementation of PackNet and related self-supervised monocular depth estimation methods from Toyota Research Institute's … | 23 | 1274 | maintenance |
| openpifpaf/openpifpaf OpenPifPaf is a PyTorch library implementing Composite Fields for semantic keypoint detection and spatio-temporal association, primarily fo… | 23 | 1262 | maintenance |
| harvardnlp/im2markup A deep learning system (built on Torch) that converts images of rendered text into presentational markup such as LaTeX or HTML, using a CNN… | 32 | 1258 | maintenance |
| ranahanocka/point2mesh Point2Mesh is a PyTorch implementation of a SIGGRAPH 2020 technique that reconstructs watertight surface meshes from input point clouds by … | 32 | 1240 | maintenance |
| microapp-store/linjiashop Linjiashop is a lightweight, open-source e-commerce (mall) system built with Spring Boot and Vue.js, MIT licensed. It ships with an admin b… | 23 | 1237 | maintenance |
| NVlabs/VoxFormer Official PyTorch implementation of VoxFormer, a CVPR 2023 highlight paper presenting a sparse voxel transformer for camera-based 3D semanti… | 31 | 1208 | maintenance |
| peng-zhihui/A-Eye A super-mini AI camera development board based on the Kendryte K210 chip, with fully open-source hardware (PCB and enclosure designs) and f… | 32 | 1190 | maintenance |
| elliottwu/unsup3d Official PyTorch implementation of the CVPR 2020 (Oral, Best Paper Award) research paper 'Unsupervised Learning of Probably Symmetric Defor… | 32 | 1189 | maintenance |
| Viveckh/Veniqa Veniqa is a full-stack open-source e-commerce solution built on the MEVN stack (MongoDB, Express.js, Vue.js, Node.js), comprising API serve… | 23 | 1189 | maintenance |
| zju3dv/snake Official research code for 'Deep Snake for Real-Time Instance Segmentation' (CVPR 2020 oral), implementing a deep contour-based instance se… | 32 | 1185 | maintenance |
| facebookresearch/meshrcnn Mesh R-CNN is Facebook AI Research's official implementation of the ICCV 2019 paper, a model that detects objects in images and predicts th… | 60 | 1161 | maintenance |
| soulteary/docker-prompt-generator A Docker-based web application that uses language models to generate and expand prompts for image generation tools like MidJourney and Stab… | 30 | 1157 | maintenance |
| dmytro-anokhin/url-image URLImage is a lightweight, pure SwiftUI view that downloads and displays images from URLs, with in-memory and on-disk caching. It serves as… | 10 | 1156 | maintenance |
| Xharlie/pointnerf Point-NeRF is a research implementation of a point-based neural radiance field method (CVPR 2022 Oral) that models scenes with neural 3D po… | 32 | 1154 | maintenance |
| s045pd/DarkNet_ChineseTrading A Python-based real-time crawler that monitors Chinese-language darknet marketplaces over Tor, with automatic account registration, login, … | 10 | 1150 | maintenance |
| bubbliiiing/yolov5-pytorch A PyTorch implementation of the YOLOv5 (v5.0) object detection model with full training, prediction, and evaluation pipelines. It is design… | 23 | 1149 | maintenance |
| joe-siyuan-qiao/DetectoRS Official PyTorch implementation of DetectoRS, a state-of-the-art object detection and instance segmentation model using Recursive Feature P… | 32 | 1147 | maintenance |
| HRNet/HRNet-Facial-Landmark-Detection Official PyTorch implementation of HRNet-based facial landmark detection from the TPAMI paper 'Deep High-Resolution Representation Learning… | 32 | 1138 | maintenance |
| Timthony/self_drive A self-driving RC car project based on Raspberry Pi and TensorFlow/Keras. It collects camera images while a human drives the car on a taped… | 32 | 1134 | maintenance |
| yu-takagi/StableDiffusionReconstruction Research codebase reproducing Takagi and Nishimoto's CVPR 2023 method for reconstructing images a person viewed from fMRI brain activity us… | 30 | 1127 | maintenance |
| bearpaw/pytorch-pose A PyTorch toolkit implementing a general pipeline for 2D single-human pose estimation, with training, inference, and evaluation interfaces … | 32 | 1121 | maintenance |
| irolaina/FCRN-DepthPrediction Reference implementation and pretrained models for FCRN (Deeper Depth Prediction with Fully Convolutional Residual Networks), predicting de… | 32 | 1118 | maintenance |
| fudan-zvg/SETR SETR (SEgmentation TRansformers) is the official PyTorch implementation of the CVPR 2021 / IJCV 2024 paper 'Rethinking Semantic Segmentatio… | 32 | 1108 | maintenance |
| aim-uofa/AdelaiDepth AdelaiDepth is an open-source toolbox for monocular depth prediction and 3D scene reconstruction from single images, containing research pr… | 32 | 1105 | maintenance |
| msracver/Relation-Networks-for-Object-Detection Official MXNet implementation of the CVPR 2018 paper 'Relation Networks for Object Detection', which adds an attention-based relation modul… | 32 | 1104 | maintenance |
| PaddlePaddle/Paddle.js Paddle.js is a browser-based deep learning inference engine for Baidu PaddlePaddle models, running via WebGL, WebGPU, or WebAssembly backen… | 23 | 1104 | maintenance |
| flibitijibibo/SDL2-CS SDL2# is a C# wrapper (P/Invoke bindings) for the SDL2 multimedia library and its extensions (SDL2_gfx, SDL2_image, SDL2_mixer, SDL2_ttf). … | 10 | 1103 | maintenance |
| hhaAndroid/mmdetection-mini A minimal, heavily annotated reimplementation of the mmdetection object detection framework, built from scratch for learning purposes. It m… | 32 | 1100 | maintenance |
| kazuto1011/deeplab-pytorch An unofficial PyTorch re-implementation of DeepLab v2 with a ResNet-101 backbone for semantic segmentation, supporting COCO-Stuff and PASCA… | 23 | 1100 | maintenance |
| sfzhang15/ATSS Official PyTorch implementation of ATSS (Adaptive Training Sample Selection), a CVPR 2020 Oral paper on object detection. It automatically … | 32 | 1086 | maintenance |
| GOATmessi8/ASFF A PyTorch implementation of YOLOv3 with the Adaptively Spatial Feature Fusion (ASFF) module and optional MobileNetV2 backbone for single-sh… | 32 | 1085 | maintenance |
| AstarLight/CPS-OCR-Engine A deep-learning-based OCR engine from SYSU DeepDriving Lab that recognizes 3755 printed Chinese characters (Level-1 character set) in elect… | 32 | 1083 | maintenance |
| WangRongsheng/XrayGLM XrayGLM is the first Chinese multimodal medical large language model that generates radiology report summaries from chest X-ray images, bui… | 29 | 1083 | maintenance |
| weiyithu/SurroundOcc SurroundOcc is the official PyTorch implementation of an ICCV 2023 paper predicting dense 3D volumetric occupancy from multi-camera images … | 43 | 1080 | maintenance |
| cardwing/Codes-for-Lane-Detection Reference implementations of lightweight lane detection CNNs, including the ENet-SAD model from the ICCV 2019 paper 'Learning Lightweight L… | 32 | 1075 | maintenance |
| sunset1995/DirectVoxGO DirectVoxGO (DVGO) is a PyTorch implementation of Direct Voxel Grid Optimization for fast neural radiance field (NeRF) reconstruction, repl… | 32 | 1075 | maintenance |
| NVIDIA-AI-IOT/trt_pose trt_pose is a Python library from NVIDIA for real-time human pose estimation accelerated with TensorRT, targeting NVIDIA Jetson and other N… | 23 | 1066 | maintenance |
| JiehangXie/PaddleBoBo PaddleBoBo is a Python project built on PaddlePaddle (with PaddleSpeech and PaddleGAN) that quickly generates a virtual streamer (VTuber) f… | 32 | 1063 | maintenance |
| ternaus/TernausNet TernausNet is a PyTorch implementation of the U-Net architecture with a VGG11 encoder pre-trained on ImageNet for image segmentation. It wa… | 32 | 1062 | maintenance |
| OpenBMB/VisCPM VisCPM is a family of open-source bilingual (Chinese/English) multimodal large models built on the 10B CPM-Bee language model, comprising V… | 29 | 1062 | maintenance |
| glyphr-studio/Glyphr-Studio-1 Glyphr Studio v1 is a free, web-based font editor aimed at hobbyists and typeface design beginners, offering vector glyph editing, kerning,… | 10 | 1062 | maintenance |
| zhaoweicai/cascade-rcnn A C++/Caffe implementation of Cascade R-CNN and other popular two-stage object detection frameworks such as Faster R-CNN, R-FCN, and FPN. I… | 32 | 1060 | maintenance |
| OFA-Sys/ONE-PEACE ONE-PEACE is a general multimodal representation model that jointly encodes vision, audio, and language modalities without initializing fro… | 29 | 1060 | maintenance |
| microsoft/Oscar Oscar is Microsoft's research code for object-semantics aligned cross-modal pre-training of vision-language models, with VinVL providing im… | 10 | 1053 | maintenance |
| damo-cv/TransReID Official PyTorch implementation of TransReID, an ICCV 2021 paper applying vision transformers to object re-identification. It provides trai… | 32 | 1052 | maintenance |
| microsoft/Cognitive-Samples-IntelligentKiosk A UWP sample application from Microsoft showcasing hands-free kiosk-style demos built on Azure Cognitive Services (Face, Computer Vision, T… | 10 | 1052 | maintenance |
| bravekingzhang/text2video A Python web application that converts text (e.g., novel passages) into narrated videos. It splits text into sentences, generates images vi… | 29 | 1049 | maintenance |
| YuwenXiong/py-R-FCN A Python implementation of R-FCN (Region-based Fully Convolutional Networks) for object detection, modified from the official MATLAB code a… | 32 | 1043 | maintenance |
| google-research/deeplab2 DeepLab2 is a TensorFlow library from Google Research providing a unified, state-of-the-art codebase for dense pixel labeling tasks such as… | 10 | 1037 | maintenance |
| goberoi/faceit A Python script that simplifies swapping faces in videos using the deepfakes/faceswap library, with training data sourced from YouTube vide… | 32 | 1032 | maintenance |
| edvardHua/PoseEstimationForMobile A TensorFlow-based library implementing CPM and Hourglass models with MobileNetV2 inverted residual modules for real-time single-person hum… | 32 | 1024 | maintenance |
| lmb-freiburg/flownet2 A Caffe fork implementing FlowNet 2.0, a deep CNN for optical flow estimation from image pairs, released with the CVPR 2017 paper. It inclu… | 32 | 1024 | maintenance |
| dwofk/fast-depth FastDepth is the official PyTorch implementation of the ICRA 2019 paper 'FastDepth: Fast Monocular Depth Estimation on Embedded Systems' fr… | 32 | 1024 | maintenance |
| HelloChenJinJun/NewFastFrame An Android component-based (modular) framework project in Java that integrates mainstream libraries like OkHttp, RxJava, Retrofit, Glide, G… | 32 | 1013 | maintenance |
| wpeebles/gangealing Official PyTorch implementation of GANgealing, a CVPR 2022 method that trains a Spatial Transformer to densely align images using GAN-gener… | 32 | 1012 | maintenance |
| cbh123/narrator A Python app that watches your webcam and generates David Attenborough-style narration of what it sees, using GPT vision models and ElevenL… | 69 | 4426 | experimental |
| mileyan/pseudo_lidar Research code implementing Pseudo-LiDAR, a CVPR 2019 method that converts image-based depth maps into pseudo-LiDAR point clouds for 3D obje… | 32 | 1005 | maintenance |
| bubbliiiing/yolov8-pytorch A PyTorch implementation of the YOLOv8 object detection model with training, prediction, and evaluation scripts. It supports training on cu… | 21 | 1005 | maintenance |
| google-research/magvit Official JAX implementation of MAGVIT, a masked generative video transformer from a CVPR 2023 paper by Google Research and CMU. It provides… | 10 | 1001 | maintenance |
| timelinize/timelinize Timelinize is an open-source personal archival suite that imports photos, videos, messages, location history, social media, and contacts fr… | 66 | 3649 | experimental |
| Everlyn-Labs/Everlyn-1 Everlyn-1 is an open autoregressive foundational video AI model from Everlyn Labs, accompanied by research on video compression/tokenizatio… | 22 | 2892 | experimental |