function: image-processing
4273 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| towhee-io/towhee Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s… | 23 | 3453 | maintenance |
| agiwhitelist/auteur auteur is an Agent Skill (SKILL.md plus recipes and scripts) for Claude Code and other AI coding agents that directs cinematic website buil… | 77 | 1017 | active |
| thomwolf/Magic-Sand Magic-Sand is a C++ openFrameworks application that operates an augmented reality sandbox by pairing a Kinect depth sensor with a projector… | 23 | 1016 | active |
| SunOner/sunone_aimbot An AI-powered aimbot for first-person shooter games that uses YOLO object detection models (YOLOv8/v10/v12) with TensorRT/ONNX acceleration… | 65 | 1014 | active |
| EvolvingLMMs-Lab/Otter Otter is a multi-modal vision-language model built on OpenFlamingo, instruction-tuned on the MIMIC-IT dataset with image and video understa… | 21 | 3438 | maintenance |
| hezarai/hezar Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta… | 77 | 1013 | active |
| GitLqr/LQRWeChat An open-source Android app that closely replicates WeChat 6.5.7, built on the RongCloud IM SDK with RxJava, Retrofit, MVP, and Glide. It su… | 72 | 3426 | maintenance |
| google-research/inksight InkSight is a Google Research system that converts photos of offline handwritten text into digital ink strokes using a ViT and mT5 encoder-… | 67 | 1008 | active |
| gausian-AI/Gausian_native_editor Gausian is a native desktop video editor built in Rust with GPU-accelerated preview (WGPU), timeline editing, and hardware decoding via Vid… | 53 | 1006 | active |
| RIFE RIFE is a deep learning model for real-time video frame interpolation, estimating intermediate flow between frames to generate smooth slow-… | 77 | 1004 | active |
| TinyLLaVA/TinyLLaVA_Factory TinyLLaVA Factory is an open-source modular PyTorch/HuggingFace codebase for training small-scale large multimodal models (LMMs) that combi… | 67 | 1004 | active |
| ZJUI-AI4H/Hulu-Med Hulu-Med is a family of open-source transparent generalist medical vision-language models ranging from 4B to 235B parameters, covering text… | 61 | 1001 | active |
| siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro… | 39 | 1001 | active |
| catalyst-team/catalyst Catalyst is a high-level PyTorch framework for deep learning research and development, focused on reproducibility, rapid experimentation, a… | 63 | 3382 | maintenance |
| hamuchiwa/AutoRCCar An open-source project that turns a hobby RC car into an autonomous self-driving vehicle using a Raspberry Pi, Arduino, camera, and ultraso… | 32 | 3342 | maintenance |
| HRNet/HRNet-Semantic-Segmentation Official PyTorch implementation of HRNet (High-Resolution Network) and the Segmentation Transformer (OCR) approach for semantic segmentatio… | 32 | 3330 | maintenance |
| JIA-Lab-research/MGM Official PyTorch implementation of Mini-Gemini, a multimodal vision-language model framework built on LLaVA that supports dense and MoE LLM… | 25 | 3326 | maintenance |
| NVIDIA/flownet2-pytorch A PyTorch implementation of FlowNet 2.0 for deep-learning-based optical flow estimation, released by NVIDIA. It provides multiple network a… | 66 | 3292 | maintenance |
| tinyvision/DAMO-YOLO DAMO-YOLO is a fast and accurate object detection framework built on PyTorch, featuring NAS-searched backbones, RepGFPN, a lightweight Zero… | 32 | 3187 | maintenance |
| tusen-ai/simpledet SimpleDet is a Python framework built on MXNet for object detection and instance recognition. It provides state-of-the-art detection models… | 32 | 3085 | maintenance |
| argman/EAST A TensorFlow re-implementation of the EAST (Efficient and Accurate Scene Text Detector) deep learning model for detecting text in natural s… | 32 | 3059 | maintenance |
| rinongal/textual_inversion Official implementation of the Textual Inversion paper, which learns new word embeddings in a frozen text-to-image (Latent Diffusion) model… | 32 | 3055 | maintenance |
| avinashpaliwal/Super-SloMo A PyTorch implementation of the Super SloMo paper for high-quality video frame interpolation, generating multiple intermediate frames to co… | 10 | 3024 | maintenance |
| libffcv/ffcv FFCV is a fast data loading system for PyTorch that dramatically increases data throughput in model training by replacing standard data loa… | 23 | 2992 | maintenance |
| NVIDIA/MinkowskiEngine Minkowski Engine is an auto-differentiation neural network library for high-dimensional sparse tensors, built on PyTorch with CUDA accelera… | 23 | 2958 | maintenance |
| xiaofengShi/CHINESE-OCR An end-to-end Chinese scene-text OCR pipeline combining CTPN for text detection, a VGG16-based orientation classifier, and CRNN with CTC fo… | 76 | 2955 | maintenance |
| fex-team/fis FIS is a front-end integrated solution and build tool from Baidu's FEX team that handles asset compilation, dependency analysis, compressio… | 32 | 2929 | maintenance |
| zhaipro/easy12306 A Python project that uses deep learning models to automatically recognize 12306 (China Railway) captchas, identifying both the Chinese tex… | 23 | 2910 | maintenance |
| nickliqian/cnn_captcha A Python project that uses convolutional neural networks built with TensorFlow to recognize character-based image captchas. It packages val… | 32 | 2880 | maintenance |
| kpzhang93/MTCNN_face_detection_alignment Reference MATLAB/Caffe implementation of MTCNN, a multi-task cascaded convolutional neural network for joint face detection and facial land… | 32 | 2862 | maintenance |
| roytseng-tw/Detectron.pytorch A PyTorch reimplementation of Facebook's Detectron object detection framework, supporting Mask R-CNN, keypoint/pose estimation, and instanc… | 10 | 2808 | maintenance |
| mahyarnajibi/SNIPER SNIPER is an efficient multi-scale training algorithm for object detection and instance segmentation that processes only context regions (c… | 32 | 2690 | maintenance |
| DominicBreuker/stego-toolkit A Docker image bundling many popular steganography tools and screening scripts for solving CTF challenges. It provides CLI scripts like che… | 32 | 2689 | maintenance |
| OFA-Sys/OFA OFA is a unified sequence-to-sequence pretrained model supporting English and Chinese that unifies cross-modality, vision, and language tas… | 32 | 2557 | maintenance |
| zzh8829/yolov3-tf2 A clean implementation of YOLOv3 and YOLOv3-tiny object detection in TensorFlow 2.0, with pre-trained Darknet weight conversion, inference,… | 32 | 2513 | maintenance |
| meijieru/crnn.pytorch A PyTorch implementation of the Convolutional Recurrent Neural Network (CRNN) for scene text recognition, based on the 2016 paper by Shi et… | 32 | 2494 | maintenance |
| xingyizhou/CenterTrack CenterTrack is a deep learning model and research codebase that performs simultaneous multi-object detection and tracking using center poin… | 32 | 2477 | maintenance |
| Zhongdao/Towards-Realtime-MOT A PyTorch codebase for the Joint Detection and Embedding (JDE) model, a fast multiple-object tracker that learns object detection and appea… | 32 | 2444 | maintenance |
| nbei/Deep-Flow-Guided-Video-Inpainting A PyTorch implementation of the CVPR 2019 paper 'Deep Flow-Guided Video Inpainting', which fills missing regions in videos by completing op… | 32 | 2376 | maintenance |
| princeton-vl/CornerNet Official research code for CornerNet, an object detection model that detects objects as paired keypoints, reproducing results from the ECCV… | 32 | 2369 | maintenance |
| smallcorgi/Faster-RCNN_TF A TensorFlow implementation of Faster R-CNN, a convolutional neural network for object detection with a region proposal network. It include… | 32 | 2341 | maintenance |
| veandco/go-sdl2 Go bindings (via cgo) for the SDL2 multimedia library, including optional SDL2_image, mixer, ttf, and gfx. It lets Go programs create windo… | 26 | 2327 | maintenance |
| nanohop/sketch-to-react-native A CLI tool that converts Sketch design files (exported as SVG) into React Native components using a deep neural network and headless Chrome… | 32 | 2313 | maintenance |
| hzwer/ICCV2019-LearningToPaint A PyTorch research implementation of the ICCV 2019 paper 'Learning to Paint With Model-based Deep Reinforcement Learning'. It trains agents… | 41 | 2308 | maintenance |
| qianqianwang68/omnimotion OmniMotion is a PyTorch implementation of the ICCV 2023 paper 'Tracking Everything Everywhere All at Once', which tracks every point in a v… | 29 | 2268 | maintenance |
| mrharicot/monodepth A TensorFlow implementation of unsupervised monocular depth estimation from single images using convolutional neural networks, based on the… | 32 | 2264 | maintenance |
| ShoufaChen/DiffusionDet PyTorch implementation of DiffusionDet, the first diffusion-model-based object detection framework (ICCV 2023 Best Paper Finalist). It prov… | 23 | 2256 | maintenance |
| hunglc007/tensorflow-yolov4-tflite A TensorFlow 2.x implementation of YOLOv4, YOLOv4-tiny, YOLOv3, and YOLOv3-tiny object detection models, with scripts that convert original… | 32 | 2253 | maintenance |
| Daniil-Osokin/lightweight-human-pose-estimation.pytorch A PyTorch implementation of Lightweight OpenPose for real-time 2D multi-person human pose estimation on CPU. It detects up to 18 body keypo… | 32 | 2241 | maintenance |
| zuoqing1988/ZQCNN ZQCNN is a lightweight deep learning inference framework written in C/C++ that runs on Windows, Linux, and ARM-Linux. It ships with demos f… | 61 | 2214 | maintenance |
| google-research/uda Google Research's reference implementation of Unsupervised Data Augmentation (UDA), a semi-supervised learning method that uses advanced da… | 10 | 2205 | maintenance |
| githubharald/SimpleHTR A Handwritten Text Recognition (HTR) system implemented in TensorFlow that recognizes text from images of single words or text lines, train… | 72 | 2183 | maintenance |
| magenta/magenta-js Magenta.js is a collection of TypeScript libraries for running inference with pre-trained Magenta machine learning models directly in the b… | 72 | 2125 | maintenance |
| AIZOOTech/FaceMaskDetection An open-source face mask detection project providing a lightweight SSD-based model (1.01M parameters) with inference code for PyTorch, Tens… | 32 | 2124 | maintenance |
| bubbliiiing/yolo3-pytorch A PyTorch implementation of the YOLOv3 object detection model with full training, prediction, and evaluation scripts. It supports training … | 23 | 2113 | maintenance |
| bgshih/crnn An implementation of the Convolutional Recurrent Neural Network (CRNN), combining CNN, RNN, and CTC loss for image-based sequence recogniti… | 32 | 2105 | maintenance |
| cymcsg/UltimateAndroid UltimateAndroid is a rapid development framework for Android apps that bundles view injection, ORM, asynchronous networking, image loading,… | 32 | 2101 | maintenance |
| tianqiraf/DouZero_For_HappyDouDiZhu A Python desktop application that applies the DouZero reinforcement-learning Dou Dizhu (Chinese card game) AI to the popular Happy DouDiZhu… | 23 | 2068 | maintenance |
| darglein/ADOP ADOP is a point-based differentiable neural rendering pipeline for scene refinement and novel view synthesis, implemented in C++/CUDA with … | 23 | 2027 | maintenance |
| Zz-ww/SadTalker-Video-Lip-Sync A Python tool built on SadTalker that generates lip-synced video from an audio file and a source video, with configurable face/lip region e… | 30 | 2010 | maintenance |
| WongKinYiu/yolor PyTorch implementation of the YOLOR paper 'You Only Learn One Representation: Unified Network for Multiple Tasks', a real-time object detec… | 23 | 2003 | maintenance |
| vaguileradiaz/tinfoleak tinfoleak is an open-source Python tool for OSINT/SOCMINT analysis of Twitter accounts, extracting structured intelligence such as user act… | 32 | 1981 | maintenance |
| onblog/BlogHelper BlogHelper is an Electron-based system tray application that helps Chinese writers publish local articles to mainstream blog platforms like… | 77 | 1967 | maintenance |
| hukaixuan19970627/yolov5_obb A PyTorch implementation of YOLOv5 extended for oriented (rotated) object detection using Circular Smooth Label (CSL) angle encoding. It pr… | 32 | 1947 | maintenance |
| alyssaxuu/animockup Animockup is a browser-based design tool for creating animated device mockups and product teaser videos. Users can place videos or images i… | 32 | 1925 | maintenance |
| PandaOCR PandaOCR is a free Windows desktop OCR tool that captures screen regions and recognizes text using many cloud OCR engines (Sogou, Tencent, … | 80 | 1921 | maintenance |
| zjhuang22/maskscoring_rcnn Official PyTorch implementation of Mask Scoring R-CNN (CVPR 2019), built on maskrcnn-benchmark. It adds a network block that learns the qua… | 32 | 1895 | maintenance |
| OkGoDoIt/OpenAI-API-dotnet An unofficial C#/.NET SDK wrapping the OpenAI API, covering chat completions (GPT-3.5/4), DALL-E image generation, embeddings, moderation, … | 23 | 1894 | maintenance |
| pierluigiferrari/ssd_keras A Keras implementation of the Single Shot MultiBox Detector (SSD) object detection architecture, with ports of the original trained weights… | 23 | 1869 | maintenance |
| magicleap/Atlas Atlas is a PyTorch-based deep learning model from Magic Leap that performs end-to-end 3D scene reconstruction from posed RGB images, produc… | 32 | 1859 | maintenance |
| bubbliiiing/faster-rcnn-pytorch A PyTorch implementation of the Faster R-CNN two-stage object detection model, supporting training on VOC-format datasets with ResNet or VG… | 23 | 1832 | maintenance |
| minivision-ai/Silent-Face-Anti-Spoofing An open-source silent face anti-spoofing (liveness detection) project by MiniVision, providing model training code, data preprocessing, tes… | 32 | 1805 | maintenance |
| dreamoving/dreamoving-project DreaMoving is the official implementation of a diffusion-based controllable video generation framework from Alibaba that produces high-qual… | 26 | 1790 | maintenance |
| princeton-vl/CornerNet-Lite CornerNet-Lite is the official PyTorch implementation of the paper 'CornerNet-Lite: Efficient Keypoint Based Object Detection', providing t… | 32 | 1771 | maintenance |
| CSAILVision/gandissect GANDissect is a PyTorch-based toolkit for visualizing and understanding the internal neurons of generative adversarial networks, showing ho… | 32 | 1764 | maintenance |
| facebookresearch/votenet VoteNet is the official PyTorch implementation of the ICCV 2019 paper 'Deep Hough Voting for 3D Object Detection in Point Clouds'. It provi… | 10 | 1761 | maintenance |
| salesforce/ALBEF Official PyTorch implementation of ALBEF, a vision-and-language pre-training method that aligns image and text representations before fusin… | 10 | 1754 | maintenance |
| iqiqiya/iqiqiya-API A PHP-based collection of free web API endpoints for parsing media from Chinese platforms (Douyin, Kuaishou, Bilibili, NetEase Music, Ximal… | 10 | 1738 | maintenance |
| experiencor/keras-yolo2 A Keras/TensorFlow implementation of the YOLOv2 real-time object detection model with support for training on custom datasets. It offers mu… | 23 | 1733 | maintenance |
| DeepMotionEditing/deep-motion-editing A PyTorch-based end-to-end library for editing and rendering 3D character motion, built on SIGGRAPH 2020 research. It provides deep learnin… | 32 | 1722 | maintenance |
| microsoft/i-Code Microsoft's i-Code is a collection of research models and frameworks for integrative, composable multimodal AI spanning vision, language, a… | 32 | 1704 | maintenance |
| Lam1360/YOLOv3-model-pruning A PyTorch implementation of YOLOv3 channel pruning (network slimming) applied to hand detection on the Oxford Hand dataset. It provides spa… | 32 | 1676 | maintenance |
| argusswift/YOLOv4-pytorch A PyTorch re-implementation of YOLOv4 object detection with variants including attentive YOLOv4 (SEnet, CBAM, CoordAttention) and MobileNet… | 23 | 1676 | maintenance |
| charlesq34/frustum-pointnets Official TensorFlow code release for the CVPR 2018 paper 'Frustum PointNets for 3D Object Detection from RGB-D Data' by Stanford and Nuro r… | 32 | 1668 | maintenance |
| autonomousvision/occupancy_networks Official PyTorch implementation of the CVPR 2019 paper 'Occupancy Networks: Learning 3D Reconstruction in Function Space'. It learns contin… | 32 | 1664 | maintenance |
| HonglinChu/SiamTrackers A PyTorch collection of Siamese-based visual object tracking models including SiamFC, SiamRPN++, SiamMask, Ocean, LightTrack, and the light… | 32 | 1652 | maintenance |
| invictus717/MetaTransformer Meta-Transformer is a research framework for unified multimodal learning that maps inputs from 12 modalities (text, images, point clouds, a… | 19 | 1647 | maintenance |
| facebookresearch/consistent_depth A research library from Facebook AI Research implementing Consistent Video Depth Estimation (SIGGRAPH 2020). It reconstructs dense, flicker… | 10 | 1634 | maintenance |
| DT42/BerryNet BerryNet is a deep learning gateway that turns edge devices like Raspberry Pi into intelligent, offline AI hubs for analyzing camera images… | 23 | 1610 | maintenance |
| experiencor/keras-yolo3 A Keras/TensorFlow implementation of YOLOv3 for object detection, supporting detection with pretrained weights, custom model training with … | 32 | 1608 | maintenance |
| ialhashim/DenseDepth Official Keras/TensorFlow implementation (with experimental PyTorch and TF2 code) of the DenseDepth paper for high-quality monocular depth … | 32 | 1606 | maintenance |
| iBase4J/iBase4J iBase4J is a Java distributed system architecture scaffold built on Spring Boot, Spring MVC, MyBatis-Plus, Dubbo/Motan, Redis, Shiro, and Q… | 32 | 1590 | maintenance |
| lufficc/SSD A high-quality, fast, modular reference implementation of the SSD (Single Shot MultiBox Detector) object detection model in PyTorch. It sup… | 23 | 1584 | maintenance |
| msracver/FCIS FCIS is the official MXNet implementation of the CVPR 2017 paper 'Fully Convolutional Instance-aware Semantic Segmentation', which won firs… | 32 | 1561 | maintenance |
| V2AI/Det3D Det3D is a PyTorch-based toolbox for 3D object detection from point clouds, offering implementations of models like PointPillars, SECOND, a… | 32 | 1561 | maintenance |
| vt-vl-lab/FGVC FGVC is a PyTorch implementation of the ECCV 2020 paper 'Flow-edge Guided Video Completion'. It completes missing regions in videos by extr… | 32 | 1553 | maintenance |
| Javacr/PyQt5-YOLOv5 A desktop GUI application built with PyQt5 that wraps YOLOv5 (v6.1) object detection models. It supports running detection on images, video… | 32 | 1547 | maintenance |
| YoYo000/MVSNet MVSNet is a deep learning architecture for depth map inference from unstructured multi-view images, and R-MVSNet is its recurrent extension… | 32 | 1546 | maintenance |
| avored/laravel-ecommerce AvoRed is an open-source headless e-commerce platform built on Laravel, exposing a GraphQL API for products, orders, carts, and user manage… | 23 | 1544 | maintenance |
| dandelin/ViLT Official PyTorch code for the ICML 2021 paper ViLT, a vision-and-language transformer that performs multimodal pre-training without convolu… | 23 | 1536 | maintenance |