function: image-processing
4273 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ingra14m/Deformable-3D-Gaussians Official PyTorch implementation of the CVPR 2024 paper 'Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction'. … | 18 | 1256 | stable |
| HisAtri/LrcApi A Flask-based API service that fetches LRC lyrics and album/artist cover art for music players, built primarily for StreamMusic and Navidro… | 70 | 1251 | active |
| Linketic/CityGaussian Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia… | 66 | 1251 | active |
| LaiFengiOS/LFLiveKit LFLiveKit is an open-source RTMP live streaming SDK for iOS written in Objective-C. It provides H264/AAC hardware encoding, GPUImage beauty… | 32 | 4391 | maintenance |
| ryokun6/ryos ryOS is a web-based desktop environment that recreates classic macOS and Windows interfaces in the browser, built with React and TypeScript… | 84 | 1244 | active |
| marian42/mesh_to_sdf A Python library that computes approximate signed distance fields (SDFs) for arbitrary triangle meshes, including non-watertight, self-inte… | 32 | 1242 | stable |
| showlab/Tune-A-Video Tune-A-Video is the official PyTorch implementation of an ICCV 2023 paper that fine-tunes pre-trained text-to-image diffusion models (like … | 31 | 4363 | maintenance |
| alibaba/Tora Tora is Alibaba's official implementation of a trajectory-oriented Diffusion Transformer (DiT) for controllable video generation, integrati… | 64 | 1241 | active |
| 0xacx/chatGPT-shell-cli A lightweight shell script that lets you chat with OpenAI's ChatGPT models and generate DALL-E images directly from the terminal, requiring… | 32 | 1241 | active |
| totaljs/framework Total.js framework is a full-featured, dependency-free web framework for Node.js written in pure JavaScript, comparable to Laravel, Django,… | 23 | 4357 | maintenance |
| KonghaYao/cn-font-split A Rust-based font subsetting tool that splits large CJK and other character fonts (otf, ttf, woff2) into small, web-ready packages with fin… | 78 | 1239 | active |
| owen2345/camaleon-cms Camaleon CMS is a dynamic and advanced content management system built on Ruby on Rails, designed as a flexible alternative to WordPress fo… | 90 | 1238 | active |
| JamesHeinrich/getID3 getID3() is a PHP library that extracts metadata and technical information from a wide range of multimedia files, including audio, video, a… | 82 | 1237 | stable |
| kellyvv/PhoneClaw PhoneClaw is a mobile-native local AI agent framework that turns phones into on-device agent runtimes, running Gemma models via LiteRT and … | 76 | 1232 | active |
| TheSmallHanCat/sora2api A self-hosted OpenAI-compatible API gateway that wraps Sora's text-to-video and image generation capabilities behind standard /v1/chat/comp… | 10 | 1232 | active |
| Tencent-Hunyuan/HunyuanCustom HunyuanCustom is a multimodal-driven customized video generation framework built on HunyuanVideo, supporting image, text, audio, and video … | 40 | 1228 | active |
| CharlyKeleb/SocialMedia-App Wooble is a fully functional social media application built with Flutter and Dart, featuring photo feeds, real-time messaging, stories, and… | 38 | 1221 | active |
| MotrixLab/SMPLer-X Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an… | 59 | 1219 | stable |
| fudan-generative-vision/champ Champ is a research framework for controllable and consistent human image animation using 3D parametric guidance (SMPL-based depth, normal,… | 25 | 4263 | maintenance |
| thClaws/thClaws thClaws is an open-source AI agent harness written in native Rust that ships as a single binary offering a desktop GUI, CLI, headless, and … | 77 | 1214 | active |
| ifzhang/FairMOT FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification… | 32 | 4245 | maintenance |
| Picsart-AI-Research/Text2Video-Zero Official implementation of Text2Video-Zero, a zero-shot text-to-video generation method that adapts text-to-image diffusion models like Sta… | 30 | 4243 | maintenance |
| LifeArchiveProject/BilibiliHistoryFetcher A Python/FastAPI backend tool that fetches, stores, and analyzes a user's Bilibili watch history, favorites, dynamics, comments, and intera… | 81 | 1205 | active |
| yanchunhuo/AutomationTest A Python-based automation testing framework supporting API automation, web UI automation (Selenium), app UI automation (Appium), and perfor… | 62 | 1205 | active |
| Artelnics/opennn OpenNN is an open-source C++ library for building, training, and deploying neural networks for advanced analytics. It is dependency-free, o… | 97 | 1199 | active |
| EvolvingLMMs-Lab/LLaVA-OneVision-2 A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis… | 72 | 1197 | active |
| Tencent-Hunyuan/HunyuanWorld-Mirror HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds… | 54 | 1195 | active |
| zai-org/VisualGLM-6B VisualGLM-6B is an open-source multimodal conversational language model supporting images, Chinese, and English, built on ChatGLM-6B with a… | 30 | 4155 | maintenance |
| goodreasonai/ScrapeServ ScrapeServ is a self-hosted API service that accepts a URL and returns the website's data along with browser screenshots, using Playwright … | 24 | 1181 | active |
| Siv3D/OpenSiv3D Siv3D (formerly OpenSiv3D) is a C++20 framework for creative coding, supporting 2D/3D games, media art, visualizers, and simulators. It pro… | 67 | 1180 | active |
| DAMO-NLP-SG/VideoLLaMA3 VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de… | 37 | 1178 | active |
| balancap/SSD-Tensorflow A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net… | 32 | 4101 | maintenance |
| DavidVentura/offline-translator An Android app that translates text, PDF/ODT documents, and images entirely offline using Firefox translation models on-device. It also off… | 88 | 1177 | active |
| Tencent-Hunyuan/MixGRPO MixGRPO is a research framework from Tencent Hunyuan implementing a mixed ODE-SDE GRPO algorithm for efficient reinforcement learning fine-… | 58 | 1177 | active |
| tjiiv-cprg/EPro-PnP EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc… | 40 | 1175 | stable |
| keinsaasforever/better-chatbot Keinsaas Navigator (formerly Better Chatbot) is an open-source, self-hostable AI chatbot workspace built with Next.js and the Vercel AI SDK… | 71 | 1172 | active |
| magicleap/SuperGluePretrainedNetwork SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t… | 32 | 4078 | maintenance |
| jenly1314/MLKit MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin… | 78 | 1171 | active |
| whyiyhw/chatgpt-wechat A self-hosted Go application that lets users safely use LLM assistants (ChatGPT, Gemini, DeepSeek, Dify workflows) inside WeChat by relayin… | 58 | 1170 | active |
| sergix44/xbackbone XBackBone is a self-hosted, lightweight file and media sharing platform with first-class ShareX support. It provides a web UI, multi-user m… | 83 | 1167 | active |
| zai-org/SCAIL-2 Official implementation of SCAIL-2, an open-source model for end-to-end controlled character animation that drives character videos from re… | 58 | 1166 | active |
| sirius-ai/LPRNet_Pytorch A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus… | 32 | 1159 | stable |
| fundamentalvision/Deformable-DETR Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl… | 32 | 4018 | maintenance |
| cloudflare/ai A monorepo of TypeScript packages and examples for building AI-powered applications on Cloudflare. It provides Vercel AI SDK and TanStack A… | 79 | 1154 | active |
| OpenGVLab/VisionLLM VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc… | 33 | 1154 | active |
| GENEXIS-AI/chromex Chromex is a Chrome MV3 side-panel extension that connects the browser to OpenAI's Codex CLI via a local native-messaging bridge. It lets u… | 72 | 1153 | active |
| Woolverine94/biniou biniou is a self-hosted web UI for 30+ generative AI models covering image, video, audio, and text generation, built with Gradio and Huggin… | 63 | 1150 | active |
| amazon-science/mm-cot Official PyTorch implementation of the paper 'Multimodal Chain-of-Thought Reasoning in Language Models', which adds vision features to a tw… | 31 | 3985 | maintenance |
| magicrew/doc7 doc7 is a Go CLI tool that converts PDFs, Office files, scans, screenshots, charts, formulas, and diagrams into AI-ready Markdown using any… | 77 | 1148 | active |
| cvg/glue-factory Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point… | 69 | 1143 | active |
| laravel/ai The Laravel AI SDK is a PHP package offering a unified, expressive API for interacting with AI providers such as OpenAI, Anthropic, and Gem… | 84 | 1140 | active |
| Anionex/agent-vision-toolkit A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, … | 79 | 1140 | active |
| HengyiWang/spann3r Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.… | 26 | 1140 | active |
| ifengzp/cocos-awesome A collection of commonly used game feature modules and shader effect implementations for the Cocos Creator game engine, written in TypeScri… | 35 | 1139 | active |
| clovaai/deep-text-recognition-benchmark Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio… | 32 | 3941 | maintenance |
| spacedeck/spacedeck-open Spacedeck Open is a free, open-source, web-based collaborative whiteboard application with rich media support, originally a commercial SaaS… | 32 | 1135 | active |
| weigert/TinyEngine A small C++ OpenGL wrapper and 2D/3D rendering engine (~2000 lines of code) that abstracts boilerplate OpenGL for windows, shaders, texture… | 23 | 1129 | active |
| geopavlakos/hamer HaMeR (Hand Mesh Recovery) is a transformer-based model that reconstructs 3D hand meshes from single monocular images using the MANO parame… | 56 | 1128 | active |
| zsyggg/paper-craft-skills A collection of Claude Code / Codex agent skills that turn academic papers into method figures, visual slide decks, and in-depth HTML artic… | 53 | 1128 | active |
| sml2h3/ddddocr-fastapi A minimal FastAPI-based REST API service wrapping the DdddOcr OCR engine, exposing endpoints for image text recognition, slide captcha matc… | 23 | 1126 | active |
| OpenGVLab/VideoMamba VideoMamba is a state space model (Mamba-based) architecture for efficient video understanding, released with code and pretrained models fr… | 25 | 1125 | active |
| FlagAI-Open/FlagAI FlagAI is a Python toolkit for training, fine-tuning, and deploying large-scale AI models across NLP, CV, and vision-language tasks. It int… | 64 | 3869 | maintenance |
| yangxue0827/RotationDetection AlphaRotate is a TensorFlow-based benchmark and toolbox for rotated (oriented) object detection, implementing detectors such as R2CNN, Reti… | 23 | 1118 | active |
| FutureUniant/Tailor Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea… | 37 | 1117 | active |
| HITsz-TMG/Uni-MoE Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,… | 68 | 1116 | active |
| THU-MIG/RepViT Official PyTorch implementation of RepViT, a family of lightweight CNNs designed by integrating efficient ViT architectural designs into Mo… | 19 | 1112 | stable |
| ATH-MaaS/Pixelle-MCP Pixelle MCP is an open-source omnimodal AIGC framework that converts ComfyUI workflows (local or RunningHub cloud) into MCP tools with zero… | 44 | 1109 | active |
| GPUOpen-LibrariesAndSDKs/Cauldron Cauldron is a C++ framework by AMD for rapid prototyping of rendering demos and samples on Vulkan and Direct3D 12. It provides glTF 2.0 loa… | 23 | 1098 | stable |
| farbrausch/fr_public An archive of Farbrausch's demoscene tools from 2001-2011, including the Werkkzeug visual content-creation tools, the V2 synthesizer, kkrun… | 32 | 3772 | maintenance |
| YOOTeam/OpenPPT OpenPPT is a web-based online presentation (PPT) editor based on ChatPPT, supporting the full workflow of creating, importing, editing, bea… | 36 | 1096 | active |
| Eyeline-Labs/Go-with-the-Flow Official implementation of the CVPR 2025 Oral paper 'Go-with-the-Flow', which controls motion in video diffusion models by replacing i.i.d.… | 42 | 1094 | active |
| yeahhe365/Gemini-Nexus Gemini Nexus is a Chrome extension (Manifest V3) that adds an AI assistant layer to the browser, integrating Gemini Web, the Gemini API, an… | 84 | 1092 | active |
| rhymes-ai/Aria Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda… | 23 | 1087 | active |
| 7thSamurai/steganography A C++ command-line tool that encrypts files with password-protected AES-256-CBC and hides them inside images using Least-Significant-Bit pi… | 32 | 1085 | stable |
| 1061700625/WeChat_Article A PyQt5 desktop application that crawls and downloads all articles from a specified WeChat official account. It uses Selenium to log in and… | 62 | 1078 | active |
| zju3dv/InfiniDepth InfiniDepth is a CVPR 2026 research library for monocular depth estimation that represents depth as neural implicit fields, allowing depth … | 52 | 1078 | active |
| symfony/ux Symfony UX is an initiative and collection of PHP/JavaScript packages that integrate frontend tools like Stimulus, Turbo, Chart.js, React, … | 98 | 1074 | active |
| orailnoor/cross-platform-llm-client PrivateLM is a cross-platform AI chat client built with Flutter that unifies local on-device LLM inference (GGUF models with Vulkan GPU acc… | 72 | 1074 | active |
| cleardusk/3DDFA A PyTorch implementation of the TPAMI 2017 paper 'Face Alignment in Full Pose Range: A 3D Total Solution' (3DDFA). It fits a 3D Morphable M… | 23 | 3677 | maintenance |
| microsoft/Biodiversity Microsoft AI for Good Lab's biodiversity research hub providing open-source AI models and tools for wildlife monitoring and conservation, i… | 88 | 1068 | active |
| gangweix/pixel-perfect-depth Pixel-Perfect Depth is a monocular depth estimation model based on pixel-space diffusion transformers that produces flying-pixel-free depth… | 49 | 1066 | active |
| open-gigaai/giga-models GigaModels is an open-source Python framework providing pipelines for training, inference, deployment, and compression of multi-modal, gene… | 61 | 1058 | active |
| YunYang1994/tensorflow-yolov3 A TensorFlow 1.x implementation of the YOLOv3 real-time object detector, reproducing the 'YOLOv3: An Incremental Improvement' paper. It sup… | 23 | 3613 | maintenance |
| henry123-boy/SpaTracker SpatialTracker is the official PyTorch implementation of a CVPR 2024 Highlight paper that tracks any 2D pixels in 3D space from RGB or RGBD… | 41 | 1057 | active |
| ShenhanQian/GaussianAvatars Official research code for GaussianAvatars, a CVPR 2024 Highlight method that creates photorealistic, fully controllable head avatars by ri… | 56 | 1054 | active |
| showlab/MotionDirector MotionDirector is a research library for customizing text-to-video diffusion models to generate videos with desired motions from a small se… | 27 | 1053 | active |
| BAAI-DCAI/Bunny Bunny is a family of lightweight multimodal vision-language models that combine plug-and-play vision encoders (EVA-CLIP, SigLIP) with langu… | 26 | 1053 | active |
| buddhi1980/mandelbulber2 Mandelbulber v2 is a cross-platform desktop application for generating and rendering photorealistic three-dimensional fractals such as Mand… | 71 | 1050 | active |
| agents-flex/agents-flex Agents-Flex is a lightweight, modular Java framework for building AI applications and agents, positioned as a Java counterpart to Spring AI… | 94 | 1046 | active |
| hanFengSan/eHunter eHunter is a Tampermonkey/userscript that injects a Vue 3-based comic reader UI into supported comic sites (EH/EXHentai, NHentai), offering… | 75 | 1045 | active |
| bs-community/blessing-skin-server Blessing Skin is a self-hosted PHP web application for uploading, managing, and sharing custom Minecraft skins and capes, restoring them fo… | 56 | 1044 | active |
| zju3dv/EfficientLoFTR Efficient LoFTR is a PyTorch implementation of a semi-dense local feature matching model that matches keypoints between image pairs with sp… | 40 | 1044 | active |
| EchoMimic EchoMimic is a series of open-source models (V1-V3) from Ant Group for audio-driven human animation, generating lifelike talking-head, port… | 50 | 1038 | active |
| liyupi/yu-picture An enterprise-grade collaborative cloud image library platform built with Vue 3, Spring Boot, Tencent COS object storage, and WebSocket. It… | 24 | 1038 | active |
| daniel-j/send2ereader A self-hostable web service for sending ebooks (EPUB, MOBI, PDF, TXT, CBZ, CBR) to a Kobo or Kindle ereader via its built-in browser using … | 39 | 1032 | active |
| jianjieyiban/JJYB_AI_VideoAutoCut JJYB_AI 智剪 is a local-first desktop AI video creation workbench that combines material analysis, smart shot segmentation, commentary script… | 70 | 1027 | active |
| vastxie/99AI 99AI is a commercially viable, self-hostable AI web platform built with Vue and Node.js that bundles AI chat, image/video/music generation,… | 35 | 1027 | active |
| aim-uofa/AdelaiDet AdelaiDet is an open-source Python toolbox built on Detectron2 that implements multiple instance-level detection and recognition algorithms… | 32 | 3477 | maintenance |
| open-mmlab/mmyolo MMYOLO is the OpenMMLab toolbox and benchmark for the YOLO series of object detection models, implemented on PyTorch. It provides unified i… | 23 | 3468 | maintenance |
| MeshAnything MeshAnything is an autoregressive transformer model that generates artist-created 3D meshes (up to 1600 faces in V2) aligned with a given s… | 31 | 1018 | active |