function: image-processing
4273 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| bilibili/Index-anisora Index-AniSora is Bilibili's open-source anime video generation model, capable of creating video shots in diverse anime styles from images, … | 62 | 2505 | active |
| InternLM/HuixiangDou HuixiangDou is an LLM-based professional knowledge assistant designed for group chat scenarios, using a three-stage pipeline of preprocess,… | 46 | 2502 | active |
| aurora-develop/aurora A Go service that exposes ChatGPT Web capabilities as an OpenAI-compatible API, including chat completions, responses, file Q&A, image gene… | 90 | 2498 | active |
| sthalles/SimCLR A PyTorch reference implementation of SimCLR, a self-supervised contrastive learning framework for learning visual representations from unl… | 23 | 2493 | stable |
| ppogg/YOLOv5-Lite YOLOv5-Lite is a lightweight object detection model family evolved from YOLOv5, with models as small as ~900KB (int8) that run 10-15+ FPS o… | 23 | 2490 | active |
| wdsqjq/FengYunWeather FengYunWeather is an open-source Android weather app written in Kotlin using MVX architecture with coroutines, OkHttp, Room, and Coil. It p… | 41 | 2456 | active |
| VadimBoev/FlappyBird A Flappy Bird clone written in pure C for Android, packaged as an APK under 100 KB using OpenGL ES 2, OpenSL ES, and Android Native Activit… | 76 | 2442 | active |
| nz-m/SocialEcho SocialEcho is a full-featured social networking platform built on the MERN stack (MongoDB, Express.js, React.js, Node.js) with automated co… | 31 | 2435 | active |
| openpaperwork/paperwork Paperwork is a personal document manager for Linux and Windows that scans, OCRs, indexes, and organizes paper documents. It provides keywor… | 10 | 2432 | active |
| wenlng/go-captcha GoCaptcha is a high-performance, modular behavioral CAPTCHA library for Go that generates interactive challenges including click, slide, dr… | 69 | 2418 | active |
| snapotter-hq/SnapOtter SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities… | 80 | 2412 | active |
| X-PLUG/mPLUG-DocOwl mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO… | 39 | 2411 | active |
| ValveResourceFormat/ValveResourceFormat Source 2 Viewer (VRF) is an open-source tool for browsing VPK archives and viewing, extracting, and decompiling Source 2 game assets such a… | 95 | 2409 | active |
| ailia-ai/ailia-models A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,… | 77 | 2389 | active |
| Cicada000/VV A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d… | 35 | 2377 | active |
| tencent-ailab/V-Express V-Express is a Python research project from Tencent AI Lab that generates talking head portrait videos from a reference image, audio, and V… | 25 | 2360 | active |
| facebookresearch/perception_models Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan… | 54 | 2355 | active |
| stephengpope/no-code-architects-toolkit A self-hostable Flask-based API that consolidates common media processing tasks—video editing, captioning, audio conversion, transcription,… | 50 | 2343 | active |
| Xilinx/PYNQ PYNQ is an open-source Python framework from AMD/Xilinx for designing embedded systems on Zynq and other adaptive computing platforms (FPGA… | 78 | 2340 | active |
| ArtalkJS/Artalk Artalk is a self-hosted, open-source comment system for blogs, websites, and web applications, with a lightweight Vanilla JS client and a G… | 85 | 2329 | active |
| Lingyan000/fluxdo FluxDO is a cross-platform third-party client for the Linux.do community forum, built with Flutter and featuring a Rust-based DNS-over-HTTP… | 78 | 2326 | active |
| getopenscreen/openscreen OpenScreen is a free, open-source desktop screen recorder and video editor for Windows, macOS, and Linux that turns raw captures into polis… | 82 | 2319 | active |
| Xiangyu-CAS/xiaohongshu-ops-skill A skill for the OpenClaw agent that turns it into a Xiaohongshu (RedNote) operations assistant, using browser automation (CDP) to analyze f… | 50 | 2317 | active |
| YvanYin/Metric3D Metric3D is the official PyTorch implementation of Metric3Dv1 and Metric3Dv2, monocular geometric foundation models that predict metric dep… | 33 | 2308 | active |
| mattiasgustavsson/libs A collection of single-file, public domain (dual MIT) C/C++ libraries covering graphics app frameworks, data structures, threading, HTTP, a… | 61 | 2300 | active |
| yawiii/ComfyUI-Prompt-Assistant A ComfyUI plugin that provides an all-in-one prompt assistant, connecting to cloud LLM/VLM APIs (Zhipu, SiliconFlow, Gemini, Baidu) and loc… | 70 | 2288 | active |
| helix-toolkit/helix-toolkit Helix Toolkit is a collection of 3D components for .NET, providing XAML/MVVM-compatible scene graphs and 3D rendering for WPF, WinUI, and A… | 84 | 2280 | active |
| kevinluosl/deepbot DeepBot is a system-level AI assistant desktop application (Electron/TypeScript) that automates enterprise and personal workflows through m… | 58 | 2276 | active |
| rushindrasinha/youtube-shorts-pipeline Verticals v3 (repo youtube-shorts-pipeline) is a Python CLI that automates producing and publishing YouTube Shorts: it researches a topic, … | 61 | 2270 | active |
| aigc-apps/EasyAnimate EasyAnimate is an end-to-end Python pipeline for high-resolution, long video and image generation based on transformer diffusion (DiT) mode… | 20 | 2270 | active |
| opendatalab/DocLayout-YOLO DocLayout-YOLO is a real-time YOLO-v10-based model for detecting document layout elements (text blocks, tables, figures, etc.) in diverse d… | 29 | 2263 | active |
| iText iText is a high-performance PDF library/SDK for Java and .NET that lets developers create, manipulate, inspect, sign, and secure PDF docume… | 93 | 2255 | stable |
| azavea/raster-vision Raster Vision is an open source Python library and low-code framework for building computer vision models on satellite, aerial, and other l… | 61 | 2242 | active |
| wxyhgk/retain-pdf RetainPDF is an open-source PDF translation tool that preserves layout, formulas, and document structure, with special support for scanned/… | 77 | 2238 | active |
| 864381832/xJavaFxTool xJavaFxTool is a cross-platform desktop application built with JavaFX that bundles dozens of small developer utilities, including encoding … | 53 | 2238 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2234 | active |
| aigc-apps/VideoX-Fun VideoX-Fun is a Python-based video generation pipeline built on Diffusion Transformer models (CogVideoX-Fun, Wan-Fun) that generates videos… | 67 | 2231 | active |
| microsoft/LLaVA-Med LLaVA-Med is a large language-and-vision assistant fine-tuned for the biomedicine domain, built on the LLaVA multimodal architecture. It su… | 40 | 2231 | active |
| NVIDIA/vid2vid A PyTorch implementation of NVIDIA's video-to-video synthesis method for generating high-resolution (e.g., 2048x1024) photorealistic videos… | 32 | 8692 | maintenance |
| alangrainger/immich-public-proxy A stateless proxy that sits in front of a self-hosted Immich instance and serves only explicitly shared photos, videos, and albums to the p… | 84 | 2224 | active |
| Hubs-Foundation/hubs Hubs is an open-source, browser-based multi-user 3D virtual world and social VR platform built with A-Frame, Three.js, and WebXR/WebRTC. It… | 82 | 2214 | active |
| nmwsharp/polyscope Polyscope is a lightweight C++/Python library and interactive viewer for 3D data such as surface meshes and point clouds. It lets you regis… | 74 | 2201 | active |
| PiranhaCMS/piranha.core Piranha CMS is a lightweight, decoupled, open-source content management system for .NET 8 built on ASP.NET Core and Entity Framework Core. … | 85 | 2192 | active |
| gridsome/gridsome Gridsome is a Vue.js-powered Jamstack framework and static site generator that builds fast, CDN-ready websites from any headless CMS, APIs,… | 32 | 8468 | maintenance |
| oil-oil/oil-motion Oil Motion is an agent-agnostic Skill that designs, generates, and integrates interactive web animations from AI-generated video. It handle… | 57 | 2156 | active |
| yyfz/Pi3 Pi3 (π³) is a feed-forward neural network for visual geometry reconstruction that eliminates the need for a fixed reference view, using a p… | 59 | 2141 | active |
| DepthAnything/Video-Depth-Anything Video Depth Anything is a transformer-based monocular depth estimation model for arbitrarily long videos, built on Depth Anything V2. It pr… | 41 | 2133 | active |
| ozgrozer/ai-renamer A Node.js CLI tool that uses local AI models (via Ollama or LM Studio) or OpenAI to intelligently rename files based on their contents, inc… | 26 | 2114 | active |
| ZiYang-xie/WorldGen WorldGen is a Python library that generates full 3D scenes in seconds from text prompts or images, supporting 360-degree consistent explora… | 53 | 2111 | active |
| BetaStreetOmnis/xhs_ai_publisher A desktop application for AI-powered content creation and automated publishing on Xiaohongshu (Rednote), built with PyQt5, FastAPI, and Pla… | 70 | 2071 | active |
| rememberber/MooTool MooTool is an all-in-one desktop developer toolbox offering dozens of handy utilities in a single GUI app, including JSON formatting, times… | 99 | 2043 | active |
| digitalsamba/claude-code-video-toolkit An AI-native video production toolkit designed for Claude Code, providing skills, commands, templates, and Python tools so an AI agent can … | 83 | 2041 | active |
| jsdecena/laracom Laracom is a free, open-source e-commerce application built on the Laravel PHP framework, providing a storefront and admin panel with produ… | 39 | 2038 | active |
| n00mkrad/flowframes Flowframes is a Windows GUI application for AI-based video frame interpolation, supporting RIFE (Pytorch & NCNN), DAIN (NCNN), and FLAVR (P… | 70 | 2033 | active |
| Kochava-Studios/witsy Witsy is a cross-platform desktop AI assistant built with Electron and Vue 3 that lets users chat with many LLM providers using their own A… | 69 | 2024 | active |
| CaliCastle/cali.so The open-source codebase for Cali Castle's personal website (cali.so), built with Next.js 16, React 19, TypeScript, and Tailwind CSS v4. It… | 67 | 2024 | active |
| cambrian-mllm/cambrian Cambrian-1 is a fully open family of vision-centric multimodal large language models (MLLMs) from NYU's VISIONx group, with training and ev… | 47 | 2013 | active |
| ppwwyyxx/wechat-dump A Python-based tool that extracts and parses WeChat message history from a rooted Android phone, decoding the local message database and me… | 54 | 2011 | active |
| 64bit/async-openai async-openai is a Rust library providing typed, async clients for the OpenAI API, covering chat completions, responses, embeddings, assista… | 93 | 2001 | active |
| Niek/chatgpt-web A single-page web interface for OpenAI-compatible chat APIs, built with Svelte, where users bring their own API key and chats are stored pr… | 75 | 1998 | active |
| cloneofsimo/lora A Python library for applying Low-Rank Adaptation (LoRA) to quickly fine-tune text-to-image diffusion models like Stable Diffusion. It prod… | 22 | 7553 | maintenance |
| adobe-research/custom-diffusion Custom Diffusion is a research codebase for efficiently fine-tuning text-to-image diffusion models like Stable Diffusion on a few example i… | 69 | 1977 | stable |
| donmccurdy/glTF-Transform glTF Transform is an SDK for reading, editing, and writing glTF 2.0 3D models in JavaScript and TypeScript, running on both Web and Node.js… | 77 | 1959 | active |
| Yuliang-Liu/Monkey Monkey is a large multi-modal model (LMM) research project from CVPR 2024 that improves image understanding via higher input resolution and… | 65 | 1951 | active |
| f0ng/captcha-killer-modified A modified version of the captcha-killer Burp Suite extension that intercepts captcha images from HTTP responses and recognizes them using … | 41 | 1949 | active |
| szczyglis-dev/py-gpt PyGPT is an open-source, all-in-one desktop AI assistant for Linux, Windows, and Mac, written in Python. It supports chat, agents, vision, … | 92 | 1903 | active |
| qqwweee/keras-yolo3 A Keras (TensorFlow backend) implementation of YOLOv3 for object detection, including Darknet weight conversion, image/video detection scri… | 32 | 7114 | maintenance |
| simonw/tools A collection of miscellaneous HTML+JavaScript single-page tools hosted at tools.simonwillison.net, almost entirely generated with LLMs as a… | 69 | 1879 | active |
| NVIDIA-AI-IOT/Lidar_AI_Solution NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil… | 72 | 1871 | active |
| MixLabPro/comfyui-mixlab-nodes A ComfyUI custom-nodes plugin that adds workflow-to-web-app conversion, screen sharing, floating video, GPT/LLM integration, 3D, speech rec… | 67 | 1861 | active |
| ofdrw/ofdrw OFDRW is an open-source Java library for reading, writing, and manipulating OFD (Open Fixed-layout Document) files, a Chinese national stan… | 93 | 1859 | active |
| ButzYung/SystemAnimatorOnline XR Animator is an AI-based full-body motion capture application that uses a single webcam with MediaPipe and TensorFlow.js to drive MMD/VRM… | 96 | 1854 | active |
| Kav-K/GPTDiscord GPTDiscord is a self-hosted Discord bot providing an all-in-one GPT interface with ChatGPT-style conversations, DALL-E image generation, AI… | 62 | 1853 | active |
| GongRzhe/Office-PowerPoint-MCP-Server A Model Context Protocol (MCP) server that lets LLM clients create, edit, and manage PowerPoint (.pptx) presentations via 32 tools built on… | 10 | 1852 | active |
| ever-co/ever-demand Ever Demand is an open-source, real-time on-demand commerce platform for building single-store shops, multi-vendor marketplaces, and on-dem… | 66 | 1848 | active |
| Tencent-Hunyuan/HunyuanVideo-I2V HunyuanVideo-I2V is Tencent's open-source image-to-video generation framework built on the HunyuanVideo diffusion model, providing PyTorch … | 54 | 1840 | active |
| xfangfang/Macast Macast is a cross-platform menu bar application that turns your computer into a DLNA Media Renderer using mpv as the playback engine. It le… | 23 | 6920 | maintenance |
| microlinkhq/browserless A Node.js library that wraps Puppeteer to provide a production-ready headless Chrome/Chromium driver with built-in screenshot, PDF generati… | 95 | 1833 | active |
| martinlaxenaire/curtainsjs curtains.js is a lightweight vanilla WebGL JavaScript library that converts HTML DOM elements containing images, videos, and canvases into … | 39 | 1824 | active |
| whotto/Video_note_generator A Python tool that converts video URLs into polished Xiaohongshu (Little Red Book) notes and blog articles. It downloads videos, transcribe… | 43 | 1812 | active |
| zju3dv/4K4D 4K4D is a research implementation of a 4D point cloud representation for real-time dynamic view synthesis at up to 4K resolution, built on … | 27 | 1806 | active |
| jd-opensource/JoyAI-Video-Edit JoyAI-Video-Edit is a real-time, instruction-guided video editing system that applies natural-language edits to live or uploaded video stre… | 57 | 1788 | active |
| syncfusion/flutter-widgets Syncfusion's Flutter widgets libraries providing high-quality UI widgets and file-format packages for building rich applications for iOS, A… | 85 | 1779 | active |
| Totoro97/NeuS Official PyTorch implementation of NeuS, a neural implicit surface reconstruction method that learns SDF-based surfaces via volume renderin… | 32 | 1776 | stable |
| githubXiaowangzi/NP-Manager NP-Manager is an Android application for APK, DEX, JAR, Smali, PDF, and media file manipulation. It provides reverse-engineering features s… | 76 | 1766 | active |
| GAP-LAB-CUHK-SZ/gaustudio GauStudio is a modular PyTorch framework for 3D Gaussian Splatting (3DGS) research and development, supporting novel view synthesis, 3D rec… | 49 | 1762 | active |
| elder-plinius/ST3GG ST3GG is an all-in-one steganography toolkit that hides secret data inside images, audio, documents, and network packets using 100+ encodin… | 63 | 1758 | active |
| NVIDIA-NeMo/Curator NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for … | 86 | 1751 | active |
| microsoft/DirectXTK12 DirectX Tool Kit for DirectX 12 is a C++ library of helper classes for writing Direct3D 12 code, covering sprite rendering, effects, textur… | 85 | 1751 | stable |
| skalesapp/skales Skales is a personal AI agent desktop and mobile application that runs locally on Windows, macOS, Linux, Android, and iOS, executing multi-… | 82 | 1745 | active |
| kijai/ComfyUI-Florence2 A ComfyUI custom node plugin that runs Microsoft's Florence-2 vision-language model for image captioning, object detection, segmentation, a… | 60 | 1743 | active |
| nolanx-ai/nolanx.ai NolanX is an open-source multi-modal agent platform for AI filmmaking that orchestrates text, image, audio, and video models into long-runn… | 52 | 1723 | active |
| fish2018/YPrompt YPrompt is a self-hosted web application that uses AI-guided conversation to elicit user requirements and automatically generate profession… | 44 | 1709 | active |
| Python3Spiders/WeiboSuperSpider A Weibo (Chinese microblog) scraping toolbox in Python covering users, topics, and comments, with extras like image downloading, sentiment … | 75 | 1705 | active |
| shubham-goel/4D-Humans 4DHumans is a Python research codebase implementing HMR 2.0, a transformer-based model for 3D human mesh recovery from single images, plus … | 59 | 1674 | active |
| Gen-Verse/MMaDA MMaDA is an open-source family of multimodal large diffusion language models that unify textual reasoning, multimodal understanding, and te… | 49 | 1671 | active |
| JiuhaiChen/BLIP3o Official implementation of the BLIP3o-Series, a unified autoregressive-plus-diffusion model for text-to-image generation and editing. It co… | 44 | 1667 | active |
| elixir-nx/bumblebee Bumblebee is an Elixir library providing pre-trained neural network models built on Axon, with integration for downloading models from Hugg… | 89 | 1666 | active |
| bytedance/Sa2VA Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understa… | 70 | 1666 | active |
| thunil/TecoGAN TecoGAN is the official source code for a temporally coherent GAN for video super-resolution, published at SIGGRAPH/ACM TOG. It includes in… | 32 | 6141 | maintenance |