function: image-processing
4273 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| pimcore/pimcore Pimcore is an open core PHP framework and platform for Product Experience Management that unifies PIM, MDM, DAM, CDP, DXP/CMS, and digital … | 95 | 3841 | active |
| google-research/scenic Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr… | 76 | 3821 | active |
| MaaEnd/MaaEnd MaaEnd is a vision-AI-powered automation assistant for the game 'Arknights: Endfield', built on MaaFramework. It captures the screen, recog… | 95 | 3762 | active |
| abhiTronix/vidgear VidGear is a high-performance, cross-platform Python framework for video processing built around multi-threaded and asynchronous pipelines.… | 76 | 3723 | active |
| facebookresearch/map-anything MapAnything is an open-source research framework from Meta and CMU for universal feed-forward metric 3D reconstruction using an end-to-end … | 77 | 3707 | active |
| ferdous-alam/GenCAD GenCAD is a research codebase for image-conditioned CAD model generation using transformer-based contrastive representations (CCIP) and dif… | 36 | 3676 | active |
| sukeesh/Jarvis Jarvis is a command-line personal assistant for Linux, macOS, and Windows written in Python. It offers 15+ task categories including weathe… | 48 | 3664 | active |
| Screenly/Anthias Anthias (formerly Screenly OSE) is a free, open-source digital signage platform that turns a Raspberry Pi or x86 PC into a networked media … | 99 | 3659 | active |
| MrNeRF/LichtFeld-Studio LichtFeld Studio is a native open-source desktop application for 3D Gaussian Splatting that combines training, real-time inspection, splat … | 92 | 3650 | active |
| EhViewer-NekoInverter/EhViewer An Android client app for browsing E-Hentai/ExHentai galleries, forked from EhViewer with a classic Material Design 2 style. It is maintain… | 84 | 3649 | active |
| NExT-GPT/NExT-GPT NExT-GPT is an end-to-end any-to-any multimodal large language model that accepts and generates arbitrary combinations of text, image, vide… | 37 | 3634 | active |
| GVCLab/PersonaLive PersonaLive is a diffusion-based framework for real-time, streamable portrait image animation, generating infinite-length expressive talkin… | 53 | 3627 | active |
| ZhaoJ9014/face.evoLVe A high-performance face recognition library built on PaddlePaddle and PyTorch, providing comprehensive tools for face-related analytics and… | 38 | 3591 | active |
| sligter/LandPPT LandPPT is an AI-powered presentation generation platform that turns a topic or uploaded documents (PDF, Word, Markdown, Excel, PPT) into p… | 81 | 3582 | active |
| Mereithhh/vanblog VanBlog is a self-hosted personal blogging system built with Next.js/TypeScript that combines a static-site-generated frontend, an admin da… | 34 | 3581 | active |
| liu-ziting/what-to-eat An AI-powered recipe generation web platform built with Vue 3 and TypeScript that creates recipes across China's eight major cuisines plus … | 45 | 3510 | active |
| mne-tools/mne-python MNE-Python is an open-source Python library for exploring, visualizing, and analyzing human neurophysiological data such as MEG, EEG, sEEG,… | 88 | 3502 | stable |
| PKU-YuanGroup/Video-LLaVA Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into… | 27 | 3500 | active |
| HeapsIO/heaps Heaps is a high-performance, cross-platform 2D and 3D game engine and graphics framework written in Haxe, created by the designer of the Ha… | 74 | 3499 | stable |
| huxingyi/dust3d Dust3D is a free, open-source, cross-platform 3D modeling application for creating low-poly 3D models quickly. It automates UV unwrapping, … | 95 | 3483 | active |
| roflcoopter/viseron Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti… | 98 | 3475 | active |
| HanaokaYuzu/Gemini-API A reverse-engineered asynchronous Python client library for Google's Gemini web app (formerly Bard), published as gemini-webapi on PyPI. It… | 92 | 3472 | active |
| huangjunsen0406/py-xiaozhi py-xiaozhi is an open-source, cross-platform multimodal AI voice assistant client written in Python, compatible with the xiaozhi-esp32 ecos… | 85 | 3463 | active |
| n3d1117/chatgpt-telegram-bot A self-hosted Telegram bot written in Python that integrates with OpenAI's official ChatGPT, DALL·E, and Whisper APIs to answer questions, … | 33 | 3463 | active |
| Grt1228/chatgpt-java An unofficial Java SDK for the OpenAI API covering all official endpoints including chat completions (GPT-3.5/GPT-4), DALL-E image generati… | 21 | 3420 | active |
| davidsandberg/facenet A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC… | 32 | 14348 | maintenance |
| paper-design/shaders Paper Shaders is a collection of zero-dependency HTML canvas/WebGL shader components for websites, available as vanilla JS and React npm pa… | 67 | 3414 | active |
| diced/zipline Zipline is a self-hosted file upload and URL shortening server with a feature-rich dashboard, designed as a ShareX-compatible upload target… | 99 | 3396 | active |
| WongKinYiu/yolov7 Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques… | 23 | 14141 | maintenance |
| timerring/bilive BILIVE is a Python application that records Bilibili live streams and danmaku 24/7, then automatically renders danmaku and AI-generated sub… | 61 | 3280 | active |
| jixiaozhong/Sonic Sonic is the official PyTorch implementation of the CVPR 2025 paper 'Sonic: Shifting Focus to Global Audio Perception in Portrait Animation… | 49 | 3274 | active |
| deepdoctection/deepdoctection deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c… | 98 | 3257 | active |
| cocktailpeanut/fluxgym FluxGym is a simple web UI for training FLUX LoRA models with low VRAM support (12GB/16GB/20GB). It combines the AI-Toolkit Gradio frontend… | 65 | 3251 | active |
| imraywang/wewrite WeWrite is a Python-based AI agent skill that automates the full WeChat Official Account content pipeline: topic selection from trending ne… | 79 | 3230 | active |
| mit-han-lab/bevfusion BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re… | 10 | 3230 | stable |
| orhun/ratty Ratty is a GPU-rendered terminal emulator written in Rust with Ratatui that supports inline 3D graphics alongside traditional 2D terminal r… | 78 | 3219 | active |
| kerlomz/captcha_trainer A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren… | 55 | 3211 | active |
| jy0205/Pyramid-Flow Pyramid Flow is the official PyTorch implementation of a training-efficient autoregressive video generation model based on pyramidal flow m… | 22 | 3211 | active |
| Pointcept Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan… | 76 | 3206 | active |
| Bionus/imgbrd-grabber Grabber is a highly customizable imageboard/booru browser and mass downloader that can fetch thousands of images from multiple booru source… | 88 | 3192 | active |
| Rudrabha/Wav2Lip Wav2Lip is the official research code for the ACM Multimedia 2020 paper 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In … | 45 | 13193 | maintenance |
| liyown/ai-trend-publish TrendPublish is a TypeScript-based automated content pipeline for WeChat Official Accounts that scrapes multiple sources (Twitter/X, RSS, s… | 82 | 3170 | active |
| Devolutions/IronRDP IronRDP is a modular Rust implementation of the Microsoft Remote Desktop Protocol (RDP), providing PDU codecs, connection/session state mac… | 97 | 3147 | active |
| darkzOGx/youtube-automation-agent AgentTube is a self-hosted Node.js application that uses AI agents to run a YouTube channel end to end: researching topics, writing scripts… | 82 | 3118 | active |
| theajack/cnchar cnchar is a comprehensive TypeScript library for Chinese character processing, offering pinyin conversion, stroke counts, stroke order draw… | 66 | 3087 | active |
| TanStack/ai TanStack AI is a type-safe, provider-agnostic TypeScript SDK for building AI applications with streaming chat, tool calling, agents, struct… | 80 | 3060 | active |
| sonos/tract Tract is Sonos' tiny, self-contained neural-network inference engine written in Rust. It loads ONNX, TensorFlow/TFLite, and NNEF models, op… | 99 | 3053 | active |
| off-grid-ai/OGAM Off Grid AI (OGAM) is a cross-platform mobile and desktop application that runs AI entirely on-device: GGUF LLM chat with vision, Whisper s… | 78 | 3045 | active |
| korlibs/korge KorGE is a modern multiplatform game engine written entirely in Kotlin, built on top of the Korlibs multimedia stack. It targets JVM/Androi… | 67 | 3042 | active |
| SharpAI/DeepCamera DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r… | 86 | 3038 | active |
| viperrcrypto/Siftly Siftly is a self-hosted, local-first web application for organizing Twitter/X bookmarks into a searchable, categorized knowledge base. It r… | 64 | 3028 | active |
| MeiGen-AI/MultiTalk MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a … | 56 | 2998 | active |
| Dicklesworthstone/llm_aided_ocr A Python tool that converts scanned PDFs to text via Tesseract OCR, then uses LLMs (local or API-based like OpenAI/Anthropic) to correct OC… | 71 | 2997 | active |
| leigest519/ScreenCoder ScreenCoder is a UI-to-code generation system that converts screenshots or design mockups into clean, editable HTML/CSS using a modular mul… | 62 | 2976 | active |
| marcotcr/lime Lime (Local Interpretable Model-agnostic Explanations) is a Python library that explains the predictions of any machine learning classifier… | 23 | 12161 | maintenance |
| mahlernim/google-timeline-visualizer An Android app (with an iPhone web app) that turns Google Maps Timeline (Location History) exports into an animated travel video. Users sel… | 83 | 2957 | active |
| sunsmarterjie/yolov12 YOLOv12 is a PyTorch implementation of attention-centric real-time object detectors, published at NeurIPS 2025. It provides detection model… | 59 | 2954 | active |
| iscyy/ultralyticsPro A PyTorch-based collection of improved YOLO-family object detection models (YOLOv5 through YOLOv13, RT-DETR) with pluggable modules for bac… | 48 | 2954 | active |
| Rust-SDL2/rust-sdl2 Rust bindings for the SDL2 multimedia library, wrapping low-level C APIs in idiomatic Rust. It provides access to graphics, audio, input, a… | 65 | 2950 | active |
| rust-headless-chrome/rust-headless-chrome A Rust library providing a high-level API to control headless Chrome or Chromium via the DevTools Protocol, serving as the Rust equivalent … | 81 | 2948 | active |
| microsoft/table-transformer Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un… | 23 | 2943 | active |
| MacPaw/OpenAI A community-maintained Swift package that wraps the OpenAI public API, supporting chat completions, responses, function calling, MCP tools,… | 94 | 2940 | active |
| openrecall/openrecall OpenRecall is an open-source, privacy-first digital memory tool that periodically takes screenshots of your screen, extracts text via local… | 43 | 2937 | active |
| sherlockchou86/VideoPipe VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates … | 54 | 2934 | active |
| jeeliz/jeelizFaceFilter A lightweight JavaScript/WebGL library for real-time face detection and tracking from a camera feed via WebRTC, designed for building augme… | 46 | 2933 | active |
| boredazfcuk/docker-icloudpd An Alpine Linux Docker container wrapping the iCloud Photos Downloader (icloudpd) utility for syncing iCloud photo libraries to a local ser… | 70 | 2932 | active |
| zeas2/Kirikiroid2 Kirikiroid2 is a cross-platform port of the Kirikiri2/KirikiriZ visual novel game engine, allowing Kirikiri-based games to run on platforms… | 23 | 2929 | active |
| InternLM/InternLM-XComposer InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u… | 38 | 2925 | active |
| JunChenMoCode/ChatGPT_JCM A Vue2 + ElementUI web management interface that aggregates OpenAI API endpoints (models, chat, images, audio, fine-tuning, files) into a g… | 22 | 2923 | active |
| Ovi/DummyJSON DummyJSON is a free hosted fake REST API that serves placeholder JSON data (products, users, carts, posts, quotes, todos, recipes) for fron… | 73 | 2911 | active |
| spliit-app/spliit Spliit is a free and open-source web application for sharing and tracking group expenses, serving as an alternative to Splitwise. It is bui… | 92 | 2907 | active |
| Saiyan-World/goku Goku is a family of flow-based (rectified flow Transformer) foundation models for joint image and video generation, released by HKU and Byt… | 23 | 2905 | active |
| TheSmallHanCat/flow2api Flow2API is a self-hosted Python/FastAPI service that exposes an OpenAI- and Gemini-compatible API on top of Google Flow (VideoFX/ImageFX) … | 60 | 2873 | active |
| nut-tree/nut.js nut.js is a cross-platform native UI automation and testing library for Node.js/TypeScript that controls mouse, keyboard, screen, and windo… | 23 | 2848 | active |
| OpenGVLab/InternImage InternImage is a large-scale CNN-based vision foundation model that uses deformable convolutions as its core operator, released with pretra… | 28 | 2846 | stable |
| Open-Cascade-SAS/OCCT Open CASCADE Technology (OCCT) is an open-source C++ development platform for 3D surface and solid modeling, CAD data exchange, and visuali… | 94 | 2831 | stable |
| youniaogu/MangaReader A cross-platform manga reading app built with React Native for Android and iOS, with tablet support. It uses a plugin-based design to aggre… | 51 | 2818 | active |
| QwenLM/Qwen-MM-Plugins A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.… | 57 | 2799 | active |
| rom1504/clip-retrieval A Python toolkit for computing CLIP embeddings for images and text and building a semantic search/retrieval system on top of them. It inclu… | 57 | 2795 | active |
| lucidrains/DALLE2-pytorch A PyTorch implementation of OpenAI's DALL-E 2 text-to-image synthesis model, focusing on the diffusion prior network that predicts image em… | 23 | 11301 | maintenance |
| OpenDCAI/Paper2Any Paper2Any is an open-source Python application that turns research papers, text, or topics into editable research figures, technical route … | 56 | 2781 | active |
| FWGS/xash3d-fwgs Xash3D FWGS is a cross-platform game engine forked from Xash3D, aimed at compatibility with the Half-Life (GoldSrc) engine while extending … | 86 | 2757 | active |
| sepinf-inc/IPED IPED is an open source digital forensic tool developed by the Brazilian Federal Police for processing and analyzing digital evidence from d… | 83 | 2737 | active |
| torinmb/mediapipe-touchdesigner A GPU-accelerated, self-contained MediaPipe plugin for TouchDesigner that runs MediaPipe vision models (face detection, face/hand/pose trac… | 86 | 2715 | active |
| bcosca/fatfree Fat-Free Framework (F3) is a lightweight PHP micro-framework condensed into a single file for building dynamic web applications quickly. It… | 84 | 2715 | stable |
| YusufB5/ASCILINE ASCILINE is a high-performance ASCII video rendering engine written in Python that converts video pixels into text-based representations. I… | 58 | 2709 | active |
| TMElyralab/MusePose MusePose is a diffusion-based, pose-guided image-to-video generation framework for creating virtual human videos, where a character in a re… | 28 | 2703 | active |
| roboflow/maestro maestro is a Python library from Roboflow that streamlines fine-tuning of multimodal vision-language models such as Florence-2, PaliGemma 2… | 62 | 2695 | active |
| chrome-php/chrome A PHP library for controlling headless Chrome/Chromium browsers via the DevTools protocol, supporting both synchronous and asynchronous usa… | 89 | 2677 | active |
| JIA-Lab-research/LISA LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati… | 31 | 2673 | active |
| Tencent/MimicMotion MimicMotion is a diffusion-based framework from Tencent for generating high-quality human motion videos guided by pose sequences, featuring… | 47 | 2653 | active |
| UniversalMediaServer/UniversalMediaServer Universal Media Server is a free, open-source DLNA, UPnP and HTTP(S) media server that streams or transcodes video, audio and images to TVs… | 97 | 2645 | active |
| ultralytics/yolov3 Ultralytics' PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection. It provides training, validation… | 67 | 10602 | maintenance |
| morethanwords/tweb Telegram Web K is the open-source TypeScript web client powering web.telegram.org/k/, based on the original Webogram and actively patched a… | 77 | 2628 | active |
| Tencent-Hunyuan/HY-World-2.0 HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from … | 58 | 2601 | active |
| UglyToad/PdfPig PdfPig is a C#/.NET library for reading and extracting text, images, annotations, forms, and metadata from PDF files, ported from Apache PD… | 91 | 2556 | active |
| jolibrain/deepdetect DeepDetect is an open-source deep learning runtime, CLI, and REST server written in C++ for training and inference across images, text, tab… | 95 | 2551 | active |
| X-PLUG/mPLUG-Owl mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and… | 36 | 2539 | active |
| BrunoLevy/geogram Geogram is a C++ programming library of geometric algorithms for geometry processing, including surface reconstruction, remeshing, Boolean … | 90 | 2527 | stable |
| kpcyrd/sn0int sn0int is a semi-automatic OSINT framework and package manager written in Rust that enumerates attack surface by processing public informat… | 60 | 2522 | active |