function: image-processing
4273 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| thtrieu/darkflow Darkflow is a Python library that translates Darknet's YOLO neural network definitions to TensorFlow, enabling real-time object detection a… | 32 | 6139 | maintenance |
| chineseocr A Python OCR toolkit that combines YOLO3-based text detection with CRNN/Dense recognition for Chinese and English text in natural scene ima… | 32 | 6123 | maintenance |
| ZJU4HealthCare/HealthGPT HealthGPT is a medical multimodal large language model family unifying medical image comprehension and generation via heterogeneous knowled… | 63 | 1655 | active |
| scito/extract_otp_secrets A Python CLI tool that extracts one-time password (TOTP/HOTP) secrets from QR codes exported by two-factor authentication apps like Google … | 93 | 1652 | active |
| XueZeyue/DanceGRPO Official implementation of DanceGRPO, a framework applying Group Relative Policy Optimization (GRPO) to fine-tune visual generation models … | 40 | 1652 | active |
| kefir500/apk-editor-studio APK Editor Studio is a free, open-source, cross-platform GUI tool for reverse-engineering Android APK files, built in C++/Qt. It lets users… | 24 | 1647 | active |
| RQLuo/MixTeX-Latex-OCR MixTeX is a multimodal OCR application that recognizes LaTeX formulas, tables, and mixed Chinese/English text from images, running entirely… | 22 | 1641 | active |
| songloft-org/songloft Songloft (formerly MiMusic) is a free, ad-free, plugin-based self-hosted music server written in Go, designed for managing music you legall… | 82 | 1636 | active |
| InterDigitalInc/CompressAI CompressAI is a PyTorch library and evaluation platform for end-to-end learned data compression research. It provides custom layers, entrop… | 73 | 1633 | active |
| yoshitomo-matsubara/torchdistill torchdistill is a modular, configuration-driven PyTorch framework for knowledge distillation and general deep learning experiments, requiri… | 86 | 1630 | active |
| yincongcyincong/MuseBot MuseBot is a self-hosted Go chatbot application that connects messaging platforms (Telegram, Discord, Slack, Lark/Feishu, DingTalk, WeCom, … | 78 | 1623 | active |
| Robbyant/lingbot-depth LingBot-Depth is a PyTorch-based model and toolkit for masked depth modeling that transforms incomplete, noisy depth sensor data into metri… | 56 | 1622 | active |
| aahnik/tgcf tgcf is a Python-based Telegram message forwarding automation tool that syncs messages between source and destination chats using either bo… | 23 | 1619 | active |
| JohnEarnest/Decker Decker is a multimedia platform and sketchpad for creating and sharing interactive documents with sound, images, hypertext, and scripted be… | 76 | 1618 | active |
| typicms/typicms TypiCMS is a modular, multilingual content management system built on the Laravel PHP framework. It ships with modules for pages, news, eve… | 95 | 1615 | active |
| Norsico/Video-Materials-AutoGEN-Workstation A self-hosted short-video production workstation that combines AI script generation (Gemini), batch TTS voiceover, AI image asset synthesis… | 54 | 1604 | active |
| meituan/YOLOv6 YOLOv6 is a single-stage object detection framework implemented in PyTorch, designed for industrial applications with a family of pretraine… | 23 | 5893 | maintenance |
| Tsuk1ko/cq-picsearcher-bot A Node.js QQ bot that performs reverse image searches via saucenao, ascii2d, soutubot.moe, and trace.moe, connecting to any OneBot 11-compa… | 74 | 1598 | active |
| Tencent-Hunyuan/HY-WorldPlay HY-WorldPlay (HY-World 1.5) is Tencent Hunyuan's open-source framework for interactive 3D world modeling, generating explorable 3D scenes f… | 55 | 1595 | active |
| Tencent-Hunyuan/HunyuanWorld-Voyager HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ… | 52 | 1591 | active |
| zju3dv/EasyVolcap EasyVolcap is a PyTorch-based library for accelerating neural volumetric video research, covering volumetric video capture, reconstruction,… | 27 | 1587 | active |
| Jamailar/Beav Beav (formerly RedBox) is a local-first AI content operations workbench for social media creators, combining a desktop app and a Chrome/Edg… | 81 | 1585 | active |
| microsoft/Mage Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage… | 57 | 1581 | active |
| Tencent/DepthCrafter DepthCrafter is a diffusion-based video depth estimation model from Tencent AI Lab that generates temporally consistent long depth sequence… | 38 | 1575 | active |
| hanmin0822/MisakaTranslator MisakaTranslator is a Windows desktop application that provides real-time machine translation for Galgames, text-based games, and manga. It… | 25 | 5756 | maintenance |
| shrimbly/node-banana Node Banana is an open-source, node-based visual workflow editor for building AI media generation pipelines. Users connect nodes on an infi… | 81 | 1560 | active |
| jpush/aurora-imui Aurora IMUI is a general-purpose instant messaging UI component library providing MessageList and InputView components, independent of any … | 23 | 5694 | maintenance |
| robocorp/rpaframework RPA Framework is a collection of open-source Python libraries and tools for Robotic Process Automation, usable from both Robot Framework an… | 99 | 1556 | active |
| kijai/ComfyUI-CogVideoXWrapper A ComfyUI custom node wrapper for CogVideoX and related video generation models (including Fun variants, CogVideoX 1.5, and Go-with-the-Flo… | 39 | 1551 | active |
| baaivision/Emu3.5 Emu3.5 is BAAI's native multimodal foundation model that jointly predicts next states across vision and language, trained on 10T+ interleav… | 43 | 1550 | active |
| hustvl/MapTR MapTR is an end-to-end transformer-based framework for online vectorized HD map construction from camera imagery in autonomous driving. It … | 27 | 1544 | active |
| facebookresearch/mmf MMF is a modular PyTorch framework for vision and language multimodal research from Facebook AI Research. It ships reference implementation… | 64 | 5633 | maintenance |
| Arthur151/ROMP ROMP is a PyTorch-based library and pip-installable API (simple-romp) for real-time monocular multi-person 3D human mesh recovery, implemen… | 23 | 1539 | stable |
| zenstory-ai/drama-skills A collection of ten agent skills for AI short-drama and comic-drama production, covering scripts, visual assets, storyboards, image/video p… | 80 | 1538 | active |
| fosslife/devtools-x DevTools-X is a cross-platform, offline-first desktop application bundling 41+ developer utilities (JSON formatting, Base64, JWT decoding, … | 72 | 1534 | active |
| WenmuZhou/PytorchOCR A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP… | 59 | 1524 | active |
| tin2tin/Pallaidium Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE… | 75 | 1523 | active |
| Tencent/TFace TFace is a research platform from Tencent Youtu Lab for trusty face analysis, covering face recognition, face security (anti-spoofing), fac… | 57 | 1523 | active |
| Taiizor/Sucrose Sucrose is a free, open-source wallpaper engine for Windows that renders interactive live wallpapers from GIFs, videos, URLs, web pages, Yo… | 92 | 1515 | active |
| ATH-MaaS/Ovis Ovis is an open-source Multimodal Large Language Model (MLLM) architecture that structurally aligns visual and textual embeddings, with rel… | 65 | 1514 | active |
| jmathai/elodie Elodie is a Python command-line tool that organizes photos, videos, and audio files based on their EXIF metadata, sorting them into folder … | 53 | 1502 | active |
| mortenjust/androidtool-mac A Mac desktop app that provides one-click screenshots, screen video recording (mp4 and gif), APK sideloading, bug reports, and custom scrip… | 23 | 5414 | maintenance |
| CUT3R/CUT3R CUT3R is the official PyTorch implementation of 'Continuous 3D Perception Model with Persistent State' (CVPR 2025 Oral), a stateful recurre… | 38 | 1489 | active |
| pur1fying/blue_archive_auto_script BAAS (Blue Archive Auto Script) is a GUI-based automation program for the mobile game Blue Archive that runs against 16:9 emulator screens.… | 72 | 1483 | active |
| Nutlope/napkins Napkins.dev is an open-source web application that turns screenshots or wireframes of website designs into working React + Tailwind code us… | 63 | 1476 | active |
| mit-han-lab/torchsparse TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a… | 26 | 1472 | active |
| sczhou/Upscale-A-Video Upscale-A-Video is a diffusion-based model for real-world video super-resolution that takes low-resolution videos and text prompts as input… | 26 | 1471 | active |
| autonomousvision/mip-splatting Mip-Splatting is a research implementation of alias-free 3D Gaussian Splatting, introducing a 3D smoothing filter and 2D Mip filter to elim… | 27 | 1468 | active |
| Habrador/Computational-geometry A C# computational geometry library for Unity implementing intersection algorithms, triangulations (Delaunay, constrained Delaunay), Vorono… | 32 | 1467 | active |
| laochiangx/Common.Utility A large collection of C# utility and helper classes covering common .NET development tasks such as Excel/CSV/PDF handling, HTTP requests, J… | 32 | 5306 | maintenance |
| bludit/bludit Bludit is a free, open-source flat-file CMS written in PHP for building websites and blogs in seconds. It stores content as JSON files, req… | 94 | 1465 | active |
| nk-o/jarallax Jarallax is a dependency-free JavaScript library that adds smooth parallax scrolling effects to background images, <img> tags, inline eleme… | 76 | 1459 | active |
| valentinfrlch/ha-llmvision LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live … | 90 | 1457 | active |
| NVlabs/Fast-FoundationStereo Fast-FoundationStereo is NVIDIA's official PyTorch implementation of a real-time zero-shot stereo matching model family, accepted to CVPR 2… | 54 | 1454 | active |
| dbolya/yolact YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++… | 51 | 5241 | maintenance |
| zylo117/Yet-Another-EfficientDet-Pytorch A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim… | 23 | 5238 | maintenance |
| tianweiy/DMD2 DMD2 is the official PyTorch implementation of Improved Distribution Matching Distillation, a NeurIPS 2024 method that distills diffusion m… | 28 | 1448 | active |
| Topdu/OpenOCR OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta… | 58 | 1440 | active |
| v-modal/vmodal_sdk_flutter VModal for Flutter is a Dart/Flutter SDK that adds multimodal video search (semantic, ASR, and OCR based) and streamed video uploads to And… | 57 | 1439 | active |
| jrzaurin/pytorch-widedeep A PyTorch library for multimodal deep learning that combines tabular data with text and images using Wide and Deep model architectures. It … | 62 | 1416 | active |
| Arthi-chaud/Meelo Meelo is a self-hosted music streaming server designed for music collectors, similar to Plex or Jellyfin but focused on music. It offers ri… | 99 | 1406 | active |
| Zejun-Yang/AniPortrait AniPortrait is a Python framework from Tencent that generates photorealistic portrait animations from an audio clip and a reference image, … | 25 | 5021 | maintenance |
| Nativ Nativ is a free, MIT-licensed macOS desktop application for running open AI models locally on Apple Silicon Macs, built on MLX-VLM. It prov… | 80 | 1397 | active |
| zju3dv/street_gaussians Street Gaussians is a research implementation of the ECCV 2024 paper 'Modeling Dynamic Urban Scenes with Gaussian Splatting', which reconst… | 40 | 1391 | active |
| Junyi42/monst3r MonST3R is the official PyTorch implementation of an ICLR 2025 paper that estimates per-timestep geometry (pointmaps) from dynamic videos i… | 36 | 1387 | active |
| arkohut/pensieve Pensieve is a privacy-focused passive screen recording application that automatically takes screenshots of your screens, indexes them with … | 85 | 1386 | active |
| gettalong/hexapdf HexaPDF is a pure Ruby library with an accompanying CLI application for creating, manipulating, merging, encrypting, signing, and optimizin… | 76 | 1382 | active |
| aws-samples/generative-ai-use-cases Generative AI Use Cases (GenU) is a well-architected sample application from AWS that demonstrates business use cases for generative AI, bu… | 92 | 1381 | active |
| yanx27/Pointnet_Pointnet2_pytorch A pure PyTorch implementation of the PointNet and PointNet++ deep learning architectures for point cloud processing. It includes training a… | 32 | 4945 | maintenance |
| autonomousvision/unimatch UniMatch is a PyTorch research library implementing a unified transformer-based model for optical flow, stereo matching, and depth estimati… | 32 | 1379 | stable |
| Meituan-AutoML/MobileVLM MobileVLM is a family of compact vision language models (1.4B-3B parameters) designed to run efficiently on mobile devices, combining small… | 17 | 1370 | active |
| PurpleDoubleD/locally-uncensored Locally Uncensored is a free, open-source desktop AI studio (built with Tauri/TypeScript) that bundles uncensored local chat, a coding agen… | 81 | 1366 | active |
| 79E/ChatGpt-Web A commercially-viable ChatGPT web application built with React and TypeScript, featuring a full admin backend for managing users, tokens, p… | 26 | 1363 | active |
| Sense-X/Co-DETR Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg… | 32 | 1360 | stable |
| ant-research/CoDeF CoDeF is the official PyTorch implementation of Content Deformation Fields, a video representation combining a canonical content field and … | 28 | 4846 | maintenance |
| pthom/imgui_bundle Dear ImGui Bundle is a batteries-included immediate-mode GUI framework for both Python and C++, built on Dear ImGui with integrated plottin… | 94 | 1353 | active |
| campsite/campsite Campsite is an open-source team communication app combining posts, chat, video calls, and AI-generated summaries for distributed teams. Thi… | 31 | 4827 | maintenance |
| huridocs/pdf-document-layout-analysis A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl… | 82 | 1349 | active |
| ParisNeo/lollms-webui LoLLMs WebUI is a local, single-user web interface for running large language models and multimodal AI systems, supporting hundreds of mode… | 71 | 4785 | maintenance |
| IrisRainbowNeko/genshin_auto_fish A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen… | 23 | 4759 | maintenance |
| LLaVA-VL/LLaVA-NeXT LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision … | 64 | 4716 | maintenance |
| open-edge-platform/geti Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru… | 98 | 1325 | active |
| yukkcat/gemini-business2api A self-hosted gateway service that exposes Gemini Business through an OpenAI-compatible API, with multi-account load balancing and an admin… | 58 | 1324 | active |
| Capsize-Games/airunner AI Runner is a privacy-focused desktop application for running local AI models offline, combining an AI chat companion with voice conversat… | 79 | 1314 | active |
| albermax/innvestigate iNNvestigate is a Python toolbox providing a common interface and out-of-the-box implementations of many neural network explanation methods… | 30 | 1308 | active |
| davidhampgonsalves/Life-Dashboard A low-power personal dashboard that repurposes a jailbroken Kindle's e-ink screen to display daily information like weather, calendar, and … | 50 | 1304 | active |
| wpilibsuite/allwpilib WPILib is the official software library suite for programming robots in the FIRST Robotics Competition (FRC) and FIRST Tech Challenge (FTC)… | 84 | 1301 | stable |
| dabit3/react-native-ai React Native AI is a full-stack framework for building cross-platform mobile AI apps with React Native and an Express server proxy. It prov… | 69 | 1297 | active |
| reg-viz/reg-suit reg-suit is a command-line tool for visual regression testing that compares current screenshots against previous snapshots and generates HT… | 79 | 1290 | active |
| TencentQQGYLab/ELLA ELLA is an Efficient Large Language Model Adapter that equips text-to-image diffusion models with LLM-based text understanding via a Timest… | 25 | 1290 | active |
| RoyalVane/CLAN Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation… | 32 | 1289 | stable |
| jordanrendric/claude-video-vision A Claude Code plugin with an MCP server that gives Claude the ability to watch and understand videos by extracting frames via ffmpeg and tr… | 58 | 1284 | active |
| AaronJackson/vrn Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses… | 32 | 4516 | maintenance |
| flutter-ml/google_ml_kit_flutter A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa… | 76 | 1278 | active |
| nianticlabs/monodepth2 Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra… | 32 | 4500 | maintenance |
| leoxiaobin/deep-high-resolution-net.pytorch Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai… | 32 | 4480 | maintenance |
| DimitarPetrov/stegify stegify is a Go command line tool and library for LSB (Least Significant Bit) steganography that hides any file inside images such as PNG a… | 23 | 1266 | stable |
| Yuliang-Liu/MonkeyOCRv2 MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2… | 58 | 1261 | active |
| MemeMeow-Studio/MemeMeow MemeMeow is a self-hosted meme/sticker management and retrieval application that lets users find images by describing the desired scene in … | 65 | 1260 | active |
| hanxi/cups-web A web-based print management application built on CUPS that turns a home USB printer into an always-available network print service. It run… | 83 | 1256 | active |