Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: image-processing

4273 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
img2threejs/img2threejs
A tool that reconstructs objects from reference images as code-only, procedural Three.js models rather than meshes or photogrammetry. It pr…
8014018active
Open3D
Open3D is an open-source C++ and Python library for 3D data processing, offering data structures, algorithms, and pipelines for point cloud…
6713913active
Cropper.js
Cropper.js is a JavaScript library for cropping images in the browser, built as customizable, extensible web components. The related 'cropp…
9113864stable
vercel/satori
Satori is a TypeScript library that converts HTML and CSS (written in JSX or React-element-like objects) into SVG strings. It handles layou…
9913857stable
python-pillow/Pillow
Pillow is the actively maintained fork of the Python Imaging Library (PIL), providing image opening, editing, and saving across many file f…
9213777stable
Curzibn/Luban
Luban 2 is an Android image compression library written in Kotlin that reverse-engineers WeChat Moments' compression strategy to produce si…
5113761active
CompVis/stable-diffusion
The original reference implementation of Stable Diffusion, a latent text-to-image diffusion model trained on LAION-5B data with a CLIP text…
3273347maintenance
lokesh/color-thief
Color Thief is a TypeScript library that extracts dominant colors and palettes from images and video in the browser and Node.js, with a sma…
9613615active
LuckSiege/PictureSelector
PictureSelector is an Android media picker library for selecting pictures, videos, and audio from the device gallery, with built-in support…
2313591stable
divamgupta/diffusionbee-stable-diffusion-ui
DiffusionBee is a free macOS desktop application that runs Stable Diffusion locally with a one-click installer and no technical setup. It p…
2313579active
microsoft/TRELLIS
TRELLIS is Microsoft's large-scale 3D asset generation model that creates high-quality 3D assets from text or image prompts. It uses a unif…
6113510active
instaloader/instaloader
Instaloader is a Python command-line tool and library for downloading pictures, videos, captions, comments, and metadata from Instagram. It…
8713240active
assimp/assimp
Open Asset Import Library (Assimp) is a C/C++ library that imports 40+ 3D model file formats (FBX, glTF, Collada, OBJ, STL, and more) into …
9313167stable
tandpfun/skill-icons
A service and icon set that renders consistent SVG skill icons for GitHub READMEs and resumes via a simple URL API (skillicons.dev). It sup…
6313019active
modelscope/DiffSynth-Studio
DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio …
7913003active
darktable-org/darktable
darktable is an open-source photography workflow application and non-destructive raw developer, acting as a virtual lighttable and darkroom…
9512992stable
lllyasviel/stable-diffusion-webui-forge
A fork/platform built on top of AUTOMATIC1111's Stable Diffusion WebUI that optimizes resource management, speeds up inference, and adds ex…
3212977active
jwagner/smartcrop.js
smartcrop.js is a JavaScript library that implements a content-aware algorithm to find good crops for images. It runs in the browser, in No…
2312955stable
gotenberg/gotenberg
Gotenberg is a Docker-based HTTP API for converting documents (HTML, URLs, Markdown, Office files) into PDFs using headless Chromium and Li…
9812942stable
ShiqiYu/libfacedetection
An open-source C++ library for CNN-based face detection in images, with the model embedded as static C source so it has no external depende…
6312784stable
octalmage/robotjs
RobotJS is a Node.js desktop automation library for controlling the mouse and keyboard and reading the screen, with native prebuilt binarie…
9312770active
google-research/vision_transformer
Google Research's official JAX/Flax implementation of Vision Transformer (ViT) and MLP-Mixer architectures, with released pretrained checkp…
7512683stable
PixEz
PixEz is a third-party Pixiv client built with Flutter that supports direct connection from mainland China without a proxy and animated ugo…
9912682active
wmjordan/PDFPatcher
PDFPatcher is a free Windows PDF toolbox built on .NET with iText and MuPDF, offering bookmark editing, page cropping/rotation, merging and…
7012633active
colmap/colmap
COLMAP is a general-purpose Structure-from-Motion (SfM) and Multi-View Stereo (MVS) pipeline for reconstructing 3D models from ordered or u…
9812564active
YaoFANGUK/video-subtitle-remover
An AI-based desktop application that removes hard-coded subtitles and text-like watermarks from videos and images using deep learning inpai…
7112553active
DayBreak-u/chineseocr_lite
An ultra-lightweight Chinese OCR toolkit combining DBNet text detection, CRNN text recognition, and an angle classifier, with total model s…
7012339active
xmu-xiaoma666/External-Attention-pytorch
A PyTorch library (fightingcv-attention) providing clean, minimal implementations of numerous attention mechanisms, MLP variants, re-parame…
6512183active
Yalantis/uCrop
uCrop is an open-source Android image cropping library by Yalantis offering flexible cropping, rotation, scaling, and compression with a bu…
4212079stable
szimek/signature_pad
Signature Pad is a zero-dependency JavaScript/TypeScript library for drawing smooth signatures on an HTML5 canvas using variable-width Bézi…
9712022stable
instantX-research/InstantID
InstantID is a tuning-free, zero-shot identity-preserving image generation method built on diffusion models, generating customized images i…
2611987active
Tongyi-MAI/Z-Image
Z-Image is a 6B-parameter text-to-image generation foundation model family built on a single-stream diffusion transformer, with a distilled…
4611944active
coil-kt/coil
Coil is a Kotlin-first image loading library for Android and Compose Multiplatform, built on Coroutines and Okio. It provides fast, lightwe…
9511881stable
qubvel-org/segmentation_models.pytorch
A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar…
7011706stable
milesial/Pytorch-UNet
A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva…
2311613active
libvips/libvips
libvips is a fast, demand-driven, horizontally threaded image processing library with low memory needs, offering around 300 operations acro…
9511602stable
flxzt/rnote
Rnote is an open-source vector-based drawing application for sketching, handwritten notes, and annotating documents and pictures, written i…
9211589active
facebookresearch/sam3
Official code for Meta's Segment Anything Model 3 (SAM 3), a unified foundation model for promptable segmentation in images and videos. It …
6311487active
kornia/kornia
Kornia is a differentiable computer vision library built on PyTorch, offering GPU-accelerated image processing, augmentations, and geometri…
8611327active
salesforce/LAVIS
LAVIS is a Python library from Salesforce AI Research providing a unified toolkit for language-vision (multimodal) intelligence, including …
6111262active
facebookresearch/dinov3
Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t…
5911249active
PointCloudLibrary/pcl
The Point Cloud Library (PCL) is a large-scale, modular open-source C++ library for 2D/3D image and point cloud processing. It provides sta…
6811101stable
imgproxy/imgproxy
imgproxy is a fast, secure standalone HTTP server written in Go (built on libvips) that resizes, processes, converts, and optimizes images …
9911030stable
ageitgey/face_recognition
A Python library and command-line tool providing a simple API for face detection, facial landmark extraction, and face recognition, built o…
6356684maintenance
google/skia
Skia is Google's open source 2D graphics library providing common APIs for drawing text, geometries, and images across hardware and softwar…
7710906stable
microsoft/TRELLIS.2
TRELLIS.2 is a 4B-parameter state-of-the-art 3D generative model from Microsoft for high-fidelity image-to-3D asset generation. It uses a f…
5710869active
cumulo-autumn/StreamDiffusion
StreamDiffusion is a Python pipeline for real-time interactive diffusion-based image generation, achieving 100+ fps on modern GPUs. It opti…
1710806active
x-hw/amazing-qr
A Python library and CLI tool (amzqr) that generates QR codes, including artistic black-and-white or colorized QR codes merged with picture…
7710805active
go-vgo/robotgo
RobotGo is a native cross-platform Go library for desktop automation, RPA, and AI computer use. It controls the mouse and keyboard, reads t…
9210781active
Automattic/node-canvas
node-canvas is a Cairo-backed implementation of the Web Canvas API for Node.js, enabling 2D drawing, text rendering, and image manipulation…
8710690active
lucidrains/denoising-diffusion-pytorch
A PyTorch implementation of Denoising Diffusion Probabilistic Models (DDPM) for generative image modeling, published on PyPI as denoising_d…
8710679active
gka/chroma.js
Chroma.js is a small, zero-dependency JavaScript library for color conversions, manipulation, and color scale generation. It supports readi…
7910574stable
amueller/word_cloud
A Python library for generating word clouds from text, with support for arbitrary masks, custom colors, and space-filling layout. It also s…
6010536stable
Acly/krita-ai-diffusion
A Krita plugin providing a streamlined interface for AI image generation, inpainting, and outpainting within the Krita painting application…
9510515active
IDEA-Research/GroundingDINO
Official PyTorch implementation of Grounding DINO, an open-set object detector that fuses a Transformer-based DINO detector with grounded v…
2110515stable
thumbor/thumbor
Thumbor is an open-source, on-demand image thumbnailing service written in Python. It crops, resizes, flips, and applies filters to images …
9010514active
esimov/caire
Caire is a content-aware image resize library written in Go, based on the seam carving algorithm. It intelligently shrinks or enlarges imag…
3110465active
easydiffusion/easydiffusion
Easy Diffusion is a 1-click installer and browser-based UI for running Stable Diffusion text-to-image generation locally on your PC. It bun…
8110455active
zyddnys/manga-image-translator
A Python tool that automatically detects, OCRs, translates, inpaints, and re-typesets text in images, primarily for manga and comics. It ru…
6510345active
lllyasviel/Fooocus
Fooocus is an offline, open-source image generation application built on Stable Diffusion XL with a Gradio interface. It simplifies text-to…
4452550maintenance
KDE/krita
Krita is a free and open-source cross-platform digital painting application built on the KDE and Qt frameworks. It provides an end-to-end s…
7710268stable
CVHub520/X-AnyLabeling
X-AnyLabeling is a cross-platform desktop application for AI-assisted annotation of text, image, video, and multimodal data. It bundles bui…
9610212active
helloianneo/ian-xiaohei-illustrations
A Codex Skill that guides AI agents to generate hand-drawn, quirky 16:9 illustrations for Chinese articles, featuring a signature 'Xiaohei'…
5910206active
Orama-Interactive/Pixelorama
Pixelorama is a free and open-source pixel art editor and multitool built with Godot, supporting sprite creation, frame-by-frame animation,…
9510195active
TencentARC/PhotoMaker
PhotoMaker is a personalized text-to-image generation method that encodes multiple reference face photos into a stacked ID embedding to gen…
2610088stable
freemocap/freemocap
FreeMoCap is a free, open-source, markerless motion capture system that uses ordinary cameras (webcams, GoPros, smartphones) to record and …
9810085active
qrohlf/trianglify
Trianglify is a JavaScript library that algorithmically generates colorful low-poly triangle mesh patterns, output as SVG images, canvas el…
4210085stable
optiscaler/OptiScaler
OptiScaler is a Windows tool that lets gamers replace upscalers in games that already support DLSS2+/FSR2+/XeSS and manage frame generation…
879987active
ImageOptim/ImageOptim
ImageOptim is a free, open-source macOS GUI application that losslessly compresses images by combining multiple optimization tools like Moz…
749969stable
open-mmlab/mmsegmentation
MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat…
239930stable
playcanvas/supersplat
SuperSplat is a free, open-source browser-based editor for inspecting, editing, optimizing, and publishing 3D Gaussian Splats. It is built …
909904active
StarTrail-org/PixelRAG
PixelRAG is a Python library and hosted service for visual retrieval-augmented generation: it renders web pages and documents into screensh…
769742active
sindresorhus/pageres
A Node.js library for capturing website screenshots at multiple resolutions using headless Chrome (Puppeteer). It supports custom CSS/JS in…
449735active
koral--/android-gif-drawable
An Android library providing Views and Drawables for displaying animated GIFs, using bundled GIFLib via JNI for efficient frame rendering. …
879647stable
CyberTimon/RapidRAW
RapidRAW is a free, open-source, non-destructive RAW photo editor and image library manager built with Rust, wgpu, React, and Tauri. It off…
819617active
mrousavy/react-native-vision-camera
A high-performance camera library for React Native offering photo/video capture, QR/barcode scanning, and JS worklet-based frame processors…
999580active
Kozea/WeasyPrint
WeasyPrint is a Python visual rendering engine for HTML and CSS that exports documents to PDF, with a CSS layout engine designed for pagina…
909531active
modelscope/facechain
FaceChain is a deep-learning toolchain from ModelScope for generating identity-preserved personal portraits (digital twins) from a single p…
309508active
PeterL1n/RobustVideoMatting
Robust Video Matting (RVM) is a deep learning model and library for real-time human video matting, using a recurrent neural network with te…
239500stable
dicebear/dicebear
DiceBear is an open source avatar library that deterministically converts seed strings (usernames, emails, IDs) into customizable SVG avata…
999421stable
PaddlePaddle/PaddleSeg
PaddleSeg is an end-to-end image segmentation toolkit built on PaddlePaddle, offering a model zoo with dozens of pre-trained models for sem…
529382active
gyroflow/gyroflow
Gyroflow is an open-source, cross-platform application that stabilizes video using gyroscope (and optionally accelerometer) motion data log…
729366stable
domlysz/BlenderGIS
A collection of Blender addons that bridge Blender with geographic data, enabling import of GIS formats like Shapefile, GeoTIFF DEM, and Op…
689326active
roboflow/rf-detr
RF-DETR is a real-time transformer-based model architecture from Roboflow for object detection, instance segmentation, and keypoint detecti…
879063active
RealSense SDK
RealSense SDK 2.0 (librealsense) is a cross-platform C++ library for Intel/RealSense depth cameras, providing depth and color streaming plu…
938974active
dusty-nv/jetson-inference
A C++/Python DNN inference library and tutorial guide ('Hello AI World') for deploying deep learning vision models on NVIDIA Jetson devices…
448969stable
infinitered/nsfwjs
NSFWJS is a JavaScript library that uses TensorFlow.js to classify images into NSFW/safety categories (Drawing, Neutral, Sexy, Hentai, Porn…
868965active
apple/ml-sharp
SHARP is a Python tool from Apple that synthesizes a photorealistic 3D Gaussian splat representation from a single photograph in under a se…
428843active
NVlabs/Sana
SANA is an efficiency-oriented PyTorch codebase for high-resolution text-to-image and text-to-video generation built on Linear Diffusion Tr…
748833active
MIC-DKFZ/nnUNet
nnU-Net is a self-configuring deep learning framework for semantic image segmentation that automatically adapts preprocessing, U-Net archit…
658829stable
carrierwaveuploader/carrierwave
CarrierWave is a Ruby gem providing a flexible, class-based solution for handling file uploads in Rack-based web applications such as Rails…
748779stable
PantsuDango/Dango-Translator
Dango-Translator (团子翻译器) is a Windows desktop application that performs real-time OCR-based translation of on-screen text ('raw' untranslat…
918751active
FoundationVision/VAR
Official PyTorch implementation of Visual Autoregressive Modeling (VAR), a NeurIPS 2024 Best Paper-winning method for scalable image genera…
488729active
hardikvasa/google-images-download
A Python command-line tool that searches and downloads hundreds of images from Google Images to local storage. It uses Selenium with Chrome…
708684active
kean/Nuke
Nuke is a Swift image loading and caching framework for Apple platforms, providing an ImagePipeline for fetching, processing, and displayin…
998656stable
MONAI
MONAI is a PyTorch-based open-source framework for deep learning in healthcare imaging, providing domain-specific transforms, 3D architectu…
878634stable
react-native-image-picker/react-native-image-picker
A React Native module that lets apps select photos and videos from the device library or capture them directly with the camera using native…
668634active
oso95/scroll-world
An agent skill (SKILL.md-compatible plugin) for Claude Code, Codex, and similar coding agents that generates immersive scroll-scrubbed 'fly…
568634active
jomjol/AI-on-the-edge-device
A firmware application for ESP32-CAM boards that uses TensorFlow Lite CNNs on-device to digitize analog utility meters (water, gas, electri…
728624active
sindresorhus/Gifski
Gifski is a macOS app wrapping the gifski encoder to convert videos into high-quality animated GIFs using pngquant's cross-frame palettes a…
938537active

← prev page 2 / 43 next →