function: image-processing
4273 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| pdfcpu/pdfcpu pdfcpu is a PDF processing library and command-line tool written in Go, supporting validation, optimization, encryption, signing, merging, … | 94 | 8816 | active |
| mfontanini/presenterm presenterm is a terminal-based slideshow tool that renders presentations written in markdown. It supports images and animated gifs, customi… | 78 | 8814 | active |
| DepthAnything/Depth-Anything-V2 Depth Anything V2 is a foundation model for monocular depth estimation, trained on 595K synthetic labeled images and 62M+ real unlabeled im… | 56 | 8757 | stable |
| fudan-generative-vision/hallo Hallo is a Python research library implementing hierarchical audio-driven visual synthesis for animating portrait images into talking-head … | 14 | 8664 | active |
| TeamWiseFlow/xiaobei Xiaobei is an open-source multi-agent system that automates social media marketing and customer acquisition for solo entrepreneurs and smal… | 90 | 8492 | active |
| Ucas-HaoranWei/GOT-OCR2.0 Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, … | 25 | 8221 | active |
| duixcom/Duix-Mobile Duix Mobile is an open-source SDK for building real-time interactive AI avatars (digital humans) that run on-device on Android, iOS, tablet… | 69 | 8211 | active |
| open-pencil/open-pencil OpenPencil is an open-source, AI-native design editor that opens and writes native Figma (.fig) files, with built-in AI chat that can creat… | 81 | 8160 | active |
| webiny/webiny-js Webiny is an open-source, self-hosted content platform built as a TypeScript framework that runs on AWS serverless infrastructure (Lambda, … | 98 | 8032 | active |
| wang-xinyu/tensorrtx A C++ collection of popular deep learning networks (YOLO variants, ResNet, MobileNet, Swin Transformer, OCR models, and more) implemented f… | 76 | 7831 | active |
| chenyme/grok2api A self-hosted multi-account API gateway for Grok Build, Grok Web, and Grok Console, written in Go with a React frontend. It exposes Grok ca… | 84 | 7594 | active |
| apple/ml-fastvlm Official implementation of FastVLM, a vision language model with an efficient hybrid vision encoder (FastViTHD) that reduces token count an… | 28 | 7415 | active |
| zai-org/GLM-OCR GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understa… | 65 | 7395 | active |
| teamchong/pxpipe pxpipe is a local TypeScript proxy that reduces Claude Code input token usage by rewriting bulky text context (system prompts, tool docs, h… | 80 | 7333 | active |
| wxWidgets/wxWidgets wxWidgets is a mature, open-source cross-platform C++ framework for building native-looking GUI desktop applications using each platform's … | 93 | 7264 | stable |
| civitai/civitai Civitai is a web platform for sharing, discovering, and discussing Stable Diffusion models, textual inversions, LoRAs, VAEs, and other gene… | 77 | 7240 | active |
| BVLC/caffe Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp… | 23 | 34553 | maintenance |
| handsomestWei/patent-disclosure-skill A Python-based agent skill for Chinese patent work that mines patentable points from project materials, drafts disclosure documents (invent… | 59 | 7193 | active |
| CMU-Perceptual-Computing-Lab/openpose OpenPose is a real-time multi-person keypoint detection library that jointly estimates human body, face, hand, and foot keypoints (135 tota… | 23 | 34423 | maintenance |
| enricoros/big-AGI Big-AGI is an open-source AI workspace web application that lets users chat with many state-of-the-art LLM providers (OpenAI, Anthropic, Ge… | 93 | 7109 | active |
| OneDragon-Anything/ZenlessZoneZero-OneDragon A Python-based automation assistant for the game Zenless Zone Zero that uses image recognition and OCR to fully automate daily tasks, dunge… | 90 | 7082 | active |
| google-research/arxiv-latex-cleaner A Python command-line tool from Google Research that cleans LaTeX source code of academic papers for arXiv submission. It removes comments,… | 80 | 7031 | active |
| apple/corenet CoreNet is Apple's deep neural network training toolkit for training standard and novel small and large-scale models, including foundation … | 45 | 7007 | active |
| zhu1090093659/dsh-web dsh-web is a plugin aggregation ecosystem package for the DeepSeek Harness (DSH) Web GUI, bundling independently installable plugins such a… | 79 | 6788 | active |
| OLMo olmOCR is an open toolkit from Ai2 that converts PDFs and image-based documents into clean, reading-order Markdown using a fine-tuned 7B vi… | 52 | 6663 | active |
| Yuliang-Liu/MonkeyOCR MonkeyOCR is a lightweight large multimodal model (LMM) for document parsing that uses a Structure-Recognition-Relation triplet paradigm to… | 60 | 6637 | active |
| vllm-project/vllm-omni vLLM-Omni is a Python framework extending vLLM for efficient inference and serving of omni-modality models, including diffusion transformer… | 83 | 6629 | active |
| TMElyralab/MuseTalk MuseTalk is a real-time, high-fidelity lip-sync model that modifies a face region in video according to input audio via latent space inpain… | 45 | 6500 | active |
| zenstory-ai/oh-story-claudecode A skill pack for Claude Code and compatible AI coding agents that provides an end-to-end workflow for writing Chinese web fiction, covering… | 80 | 6463 | active |
| mindee/doctr docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe… | 90 | 6336 | active |
| ByteDance-Seed/Depth-Anything-3 Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,… | 59 | 6268 | active |
| PaddlePaddle/PaddleX PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti… | 92 | 6255 | active |
| RangiLyu/nanodet NanoDet-Plus is a super fast, lightweight anchor-free object detection model implemented in PyTorch, with model sizes as small as 980KB (IN… | 23 | 6253 | stable |
| ldqk/Masuit.Tools Masuit.Tools is a Swiss-army utility library for C#/.NET bundling dozens of commonly needed helpers: encryption, reflection, LINQ and expre… | 96 | 6182 | active |
| ByteDance-Seed/Bagel BAGEL is an open-source unified multimodal foundation model with 7B active parameters (14B total) trained on interleaved multimodal data. I… | 55 | 6165 | active |
| basketikun/chatgpt2api A self-hosted reverse-engineered implementation of ChatGPT's official web interfaces, exposing OpenAI-compatible API endpoints for text gen… | 77 | 6130 | active |
| open-edge-platform/anomalib Anomalib is a Python deep learning library for anomaly detection, offering state-of-the-art unsupervised algorithms for detecting and local… | 98 | 6112 | active |
| Akegarasu/lora-scripts SD-Trainer is a GUI application and set of scripts for training LoRA and Dreambooth fine-tunes of Stable Diffusion diffusion models, wrappi… | 66 | 6109 | active |
| harfbuzz/harfbuzz HarfBuzz is a text shaping engine that converts Unicode text into properly positioned glyph output for any writing system, supporting OpenT… | 99 | 6051 | stable |
| bytedance/LatentSync LatentSync is an end-to-end lip-sync framework from ByteDance based on audio-conditioned latent diffusion models, using Stable Diffusion to… | 33 | 6049 | active |
| CGAL/cgal CGAL (Computational Geometry Algorithms Library) is a mature C++ library providing efficient and reliable algorithms for computational geom… | 95 | 6031 | stable |
| OpenSenseNova/SenseNova-U1 SenseNova-U is a series of open-weight unified multimodal models (e.g., SenseNova-U1.5-8B-MoT) built on the NEO-unify architecture that com… | 59 | 6017 | active |
| Doubiiu/ToonCrafter ToonCrafter is a generative model that interpolates two cartoon images into a short animation by leveraging pre-trained image-to-video diff… | 29 | 6008 | stable |
| JannisX11/blockbench Blockbench is a free, open-source desktop and web application for editing low-poly 3D models with pixel art textures. It includes modeling,… | 98 | 5877 | active |
| jindrapetrik/jpexs-decompiler JPEXS Free Flash Decompiler (FFDec) is an open-source Java application for decompiling and editing Adobe Flash SWF files. It extracts resou… | 93 | 5838 | active |
| snailyp/gemini-balance A Python FastAPI proxy and load-balancing service for the Google Gemini API that rotates multiple API keys and exposes both Gemini and Open… | 41 | 5822 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5801 | active |
| DeepLabCut/DeepLabCut DeepLabCut is an open-source Python toolbox for markerless 2D and 3D pose estimation of user-defined body parts using deep neural networks … | 89 | 5757 | stable |
| pjreddie/darknet Darknet is an open-source neural network framework written in C and CUDA, best known as the original home of the YOLO real-time object dete… | 32 | 26499 | maintenance |
| Eventual-Inc/Daft Daft is a high-performance distributed data engine with a Python dataframe API, implemented in Rust, designed for AI and multimodal workloa… | 95 | 5740 | active |
| talebook/talebook Talebook is a self-hosted personal e-book library management web application built on Calibre with a Nuxt 4 + Vue 3 frontend. It supports m… | 94 | 5732 | active |
| ramjke/Translumo Translumo is a Windows desktop application that performs real-time screen translation by capturing on-screen text with OCR and translating … | 72 | 5725 | active |
| open-mmlab/OpenPCDet OpenPCDet is a PyTorch-based open-source toolbox for LiDAR-based 3D object detection. It provides official implementations of models like P… | 53 | 5699 | active |
| apple/ml-depth-pro Depth Pro is Apple's reference implementation of a foundation model for zero-shot metric monocular depth estimation, producing sharp high-r… | 30 | 5696 | active |
| zxlie/FeHelper FeHelper is an all-in-one browser extension toolbox for Chrome, Firefox, and Edge (Manifest V3) that bundles 30+ developer utilities such a… | 76 | 5661 | active |
| ConnectAI-E/feishu-openai A self-hosted Go application that integrates OpenAI models (GPT-4, GPT-4V, DALL·E-3, Whisper) into Feishu/Lark as a chatbot. It supports vo… | 35 | 5637 | active |
| lemonade-sdk/lemonade Lemonade is a local AI server that runs optimized LLMs (plus image, speech, and embedding models) on your own GPU and NPU, exposing OpenAI-… | 82 | 5604 | active |
| sohzm/cheating-daddy Cheating Daddy is a free, open-source Electron desktop app that acts as a real-time AI assistant during video calls, interviews, and meetin… | 78 | 5589 | active |
| Vision-CAIR/MiniGPT-4 Official code for MiniGPT-4 and MiniGPT-v2, vision-language models that align a frozen visual encoder with a frozen LLM (Vicuna) via a sing… | 30 | 25619 | maintenance |
| cinder/Cinder Cinder is a free, open-source C++ library for professional-quality creative coding, covering graphics, audio, video, and computational geom… | 61 | 5538 | active |
| father-bot/chatgpt_telegram_bot A self-hostable Telegram bot that brings ChatGPT and Claude models into Telegram using your own OpenAI/Anthropic/OpenRouter API keys. It su… | 78 | 5527 | active |
| modstart-lib/aigcpanel AIGCPanel is an open-source, all-in-one AI digital human desktop application built with TypeScript, Vue3, and Electron for Windows, macOS, … | 88 | 5511 | active |
| mifi/editly Editly is a declarative non-linear video editing tool and framework built on Node.js and ffmpeg, offering both a CLI and a JavaScript API. … | 32 | 5484 | active |
| mayocream/koharu Koharu is a local-first desktop application that automates manga translation using machine learning, combining text/bubble detection, OCR, … | 82 | 5477 | active |
| isl-org/MiDaS MiDaS is a Python library with pretrained models for robust monocular depth estimation from a single image, based on the TPAMI 2022 paper a… | 10 | 5419 | stable |
| ReShade ReShade is a generic post-processing injector for games and video software that hooks into Direct3D 9-12, OpenGL, and Vulkan applications t… | 77 | 5415 | active |
| agentejo/cockpit Cockpit is a self-hosted, API-first headless CMS written in PHP that adds content management to any site via REST/JSON APIs. It offers flex… | 32 | 5394 | active |
| deepseek-ai/DeepSeek-VL2 DeepSeek-VL2 is a series of Mixture-of-Experts vision-language models (Tiny, Small, and 4.5B activated parameters) with inference code and … | 25 | 5374 | active |
| LaoFeng-mouse/flyingmouse-format FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop… | 79 | 5347 | active |
| OpenSenseNova/SenseNova-Skills A collection of modular Agent Skills for SenseNova models that add end-to-end office capabilities such as slide-deck generation, Excel data… | 59 | 5322 | active |
| mosra/magnum Magnum is a lightweight and modular C++11/C++14 graphics middleware library providing a slim wrapper over OpenGL, OpenGL ES, WebGL and Vulk… | 67 | 5204 | stable |
| NVIDIAGameWorks/kaolin Kaolin is NVIDIA's PyTorch library of GPU-optimized modules for 3D deep learning research, covering meshes, point clouds, and 3D Gaussian s… | 68 | 5164 | active |
| tixl3d/tixl TiXL (formerly Tooll3) is a free, open-source desktop application for creating realtime motion graphics, combining graph-based procedural c… | 90 | 5087 | active |
| libigl/libigl libigl is a simple C++ geometry processing library offering a wide range of algorithms for triangle and tetrahedral meshes, including defor… | 67 | 5081 | stable |
| AILab-CVC/VideoCrafter VideoCrafter is an open-source video generation and editing toolbox built on video diffusion models, offering Text-to-Video and Image-to-Vi… | 48 | 5072 | active |
| RandyGaul/cute_headers A collection of cross-platform, single-file C/C++ header-only libraries with no dependencies, primarily aimed at game development. It inclu… | 75 | 5056 | active |
| Deci-AI/super-gradients SuperGradients is an open-source PyTorch-based training library for building, training, and fine-tuning state-of-the-art computer vision mo… | 54 | 5054 | active |
| AgnesAI-Labs/AgnesAI-Models Official gateway and model catalog for Agnes AI, providing OpenAI-compatible API access to multimodal foundation models covering text, imag… | 58 | 5051 | active |
| Emby Server Emby Server is a self-hosted personal media server that organizes videos, music, photos, and live TV into a rich library and streams them t… | 23 | 4965 | active |
| ethanniser/NextFaster NextFaster is a highly performant e-commerce template built with Next.js 15, featuring partial prerendering, Server Actions, Drizzle ORM on… | 66 | 4878 | active |
| deepjavalibrary/djl Deep Java Library (DJL) is an engine-agnostic, high-level deep learning framework for Java that supports backends like PyTorch, TensorFlow,… | 80 | 4851 | active |
| sgoudelis/ground-station Ground Station is an open-source, browser-based application suite for tracking satellites and celestial targets, controlling station hardwa… | 81 | 4737 | active |
| PistonDevelopers/piston Piston is a modular game engine written in Rust, providing core libraries for windowing, event handling, and 2D graphics with a backend-agn… | 67 | 4702 | stable |
| dankamongmen/notcurses Notcurses is a C library (with C++, Rust, and Python bindings) for building complex, vibrant textual user interfaces on modern terminal emu… | 76 | 4685 | active |
| Tencent/TNN TNN is a high-performance, lightweight deep learning inference framework developed by Tencent Youtu Lab, supporting mobile, desktop, and se… | 32 | 4649 | active |
| ApostropheCMS ApostropheCMS is a full-featured, open-source content management system and framework built with Node.js and MongoDB. It combines in-contex… | 77 | 4613 | active |
| leaferjs/leafer-ui LeaferJS (leafer-ui) is a lightweight, zero-dependency TypeScript Canvas engine for building interactive 2D graphics, with a DOM-like API, … | 92 | 4397 | active |
| LibrePDF/OpenPDF OpenPDF is an open-source Java library (LGPL/MPL licensed, forked from iText 4) for creating, editing, rendering, and encrypting PDF docume… | 93 | 4360 | active |
| crmne/ruby_llm RubyLLM is a Ruby framework providing a unified, expressive interface to all major AI providers (OpenAI, Anthropic, Google, Ollama, and any… | 83 | 4336 | active |
| facebookresearch/vggt-omega VGGT-Omega is a research library from Oxford VGG and Meta AI providing pretrained transformer models for 3D vision tasks such as camera pos… | 58 | 4284 | active |
| olcPixelGameEngine olcPixelGameEngine is a single-header C++ framework for fast pixel-based 2D drawing and user input, used in javidx9's YouTube tutorials. It… | 46 | 4240 | active |
| mbloch/mapshaper Mapshaper is a JavaScript tool for editing and converting geospatial vector data formats such as Shapefile, GeoJSON, TopoJSON, GeoPackage, … | 94 | 4172 | active |
| Spark NLP Spark NLP is an open-source natural language processing library built natively on Apache Spark, providing scalable NLP annotations and tran… | 96 | 4160 | stable |
| QwenLM/Qwen3-Omni Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, an… | 52 | 3999 | active |
| area17/twill Twill is an open source CMS toolkit for Laravel that lets developers rapidly build custom, feature-rich admin consoles with pre-built Vue.j… | 95 | 3974 | active |
| ali-vilab/VACE VACE is the official implementation of an all-in-one video creation and editing model from Tongyi Lab, built on Wan2.1 diffusion models. It… | 41 | 3936 | active |
| yuka-friends/Windrecorder Windrecorder is a local-first screen recording and memory search app for Windows, an open-source alternative to Rewind.ai and Microsoft Rec… | 47 | 3930 | active |
| hardkoded/puppeteer-sharp PuppeteerSharp is a .NET port of the official Node.js Puppeteer API for controlling headless or headful Chrome and Firefox. It supports nav… | 99 | 3917 | active |
| refinery/refinerycms Refinery CMS is an open source, extendable content management system built as a Ruby on Rails engine, supporting Rails 6.1 through 8.1+ and… | 86 | 3907 | active |
| claraverse-space/ClaraVerse ClaraVerse is a self-hosted, privacy-focused AI workspace that combines chat, multi-agent teams, a visual workflow builder, RAG document pr… | 80 | 3894 | active |