Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: image-processing

4273 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
thtrieu/darkflow
Darkflow is a Python library that translates Darknet's YOLO neural network definitions to TensorFlow, enabling real-time object detection a…
326139maintenance
chineseocr
A Python OCR toolkit that combines YOLO3-based text detection with CRNN/Dense recognition for Chinese and English text in natural scene ima…
326123maintenance
ZJU4HealthCare/HealthGPT
HealthGPT is a medical multimodal large language model family unifying medical image comprehension and generation via heterogeneous knowled…
631655active
scito/extract_otp_secrets
A Python CLI tool that extracts one-time password (TOTP/HOTP) secrets from QR codes exported by two-factor authentication apps like Google …
931652active
XueZeyue/DanceGRPO
Official implementation of DanceGRPO, a framework applying Group Relative Policy Optimization (GRPO) to fine-tune visual generation models …
401652active
kefir500/apk-editor-studio
APK Editor Studio is a free, open-source, cross-platform GUI tool for reverse-engineering Android APK files, built in C++/Qt. It lets users…
241647active
RQLuo/MixTeX-Latex-OCR
MixTeX is a multimodal OCR application that recognizes LaTeX formulas, tables, and mixed Chinese/English text from images, running entirely…
221641active
songloft-org/songloft
Songloft (formerly MiMusic) is a free, ad-free, plugin-based self-hosted music server written in Go, designed for managing music you legall…
821636active
InterDigitalInc/CompressAI
CompressAI is a PyTorch library and evaluation platform for end-to-end learned data compression research. It provides custom layers, entrop…
731633active
yoshitomo-matsubara/torchdistill
torchdistill is a modular, configuration-driven PyTorch framework for knowledge distillation and general deep learning experiments, requiri…
861630active
yincongcyincong/MuseBot
MuseBot is a self-hosted Go chatbot application that connects messaging platforms (Telegram, Discord, Slack, Lark/Feishu, DingTalk, WeCom, …
781623active
Robbyant/lingbot-depth
LingBot-Depth is a PyTorch-based model and toolkit for masked depth modeling that transforms incomplete, noisy depth sensor data into metri…
561622active
aahnik/tgcf
tgcf is a Python-based Telegram message forwarding automation tool that syncs messages between source and destination chats using either bo…
231619active
JohnEarnest/Decker
Decker is a multimedia platform and sketchpad for creating and sharing interactive documents with sound, images, hypertext, and scripted be…
761618active
typicms/typicms
TypiCMS is a modular, multilingual content management system built on the Laravel PHP framework. It ships with modules for pages, news, eve…
951615active
Norsico/Video-Materials-AutoGEN-Workstation
A self-hosted short-video production workstation that combines AI script generation (Gemini), batch TTS voiceover, AI image asset synthesis…
541604active
meituan/YOLOv6
YOLOv6 is a single-stage object detection framework implemented in PyTorch, designed for industrial applications with a family of pretraine…
235893maintenance
Tsuk1ko/cq-picsearcher-bot
A Node.js QQ bot that performs reverse image searches via saucenao, ascii2d, soutubot.moe, and trace.moe, connecting to any OneBot 11-compa…
741598active
Tencent-Hunyuan/HY-WorldPlay
HY-WorldPlay (HY-World 1.5) is Tencent Hunyuan's open-source framework for interactive 3D world modeling, generating explorable 3D scenes f…
551595active
Tencent-Hunyuan/HunyuanWorld-Voyager
HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ…
521591active
zju3dv/EasyVolcap
EasyVolcap is a PyTorch-based library for accelerating neural volumetric video research, covering volumetric video capture, reconstruction,…
271587active
Jamailar/Beav
Beav (formerly RedBox) is a local-first AI content operations workbench for social media creators, combining a desktop app and a Chrome/Edg…
811585active
microsoft/Mage
Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage…
571581active
Tencent/DepthCrafter
DepthCrafter is a diffusion-based video depth estimation model from Tencent AI Lab that generates temporally consistent long depth sequence…
381575active
hanmin0822/MisakaTranslator
MisakaTranslator is a Windows desktop application that provides real-time machine translation for Galgames, text-based games, and manga. It…
255756maintenance
shrimbly/node-banana
Node Banana is an open-source, node-based visual workflow editor for building AI media generation pipelines. Users connect nodes on an infi…
811560active
jpush/aurora-imui
Aurora IMUI is a general-purpose instant messaging UI component library providing MessageList and InputView components, independent of any …
235694maintenance
robocorp/rpaframework
RPA Framework is a collection of open-source Python libraries and tools for Robotic Process Automation, usable from both Robot Framework an…
991556active
kijai/ComfyUI-CogVideoXWrapper
A ComfyUI custom node wrapper for CogVideoX and related video generation models (including Fun variants, CogVideoX 1.5, and Go-with-the-Flo…
391551active
baaivision/Emu3.5
Emu3.5 is BAAI's native multimodal foundation model that jointly predicts next states across vision and language, trained on 10T+ interleav…
431550active
hustvl/MapTR
MapTR is an end-to-end transformer-based framework for online vectorized HD map construction from camera imagery in autonomous driving. It …
271544active
facebookresearch/mmf
MMF is a modular PyTorch framework for vision and language multimodal research from Facebook AI Research. It ships reference implementation…
645633maintenance
Arthur151/ROMP
ROMP is a PyTorch-based library and pip-installable API (simple-romp) for real-time monocular multi-person 3D human mesh recovery, implemen…
231539stable
zenstory-ai/drama-skills
A collection of ten agent skills for AI short-drama and comic-drama production, covering scripts, visual assets, storyboards, image/video p…
801538active
fosslife/devtools-x
DevTools-X is a cross-platform, offline-first desktop application bundling 41+ developer utilities (JSON formatting, Base64, JWT decoding, …
721534active
WenmuZhou/PytorchOCR
A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP…
591524active
tin2tin/Pallaidium
Pallaidium is a free, open-source generative AI movie studio implemented as a Blender add-on integrated into the Video Sequence Editor (VSE…
751523active
Tencent/TFace
TFace is a research platform from Tencent Youtu Lab for trusty face analysis, covering face recognition, face security (anti-spoofing), fac…
571523active
Taiizor/Sucrose
Sucrose is a free, open-source wallpaper engine for Windows that renders interactive live wallpapers from GIFs, videos, URLs, web pages, Yo…
921515active
ATH-MaaS/Ovis
Ovis is an open-source Multimodal Large Language Model (MLLM) architecture that structurally aligns visual and textual embeddings, with rel…
651514active
jmathai/elodie
Elodie is a Python command-line tool that organizes photos, videos, and audio files based on their EXIF metadata, sorting them into folder …
531502active
mortenjust/androidtool-mac
A Mac desktop app that provides one-click screenshots, screen video recording (mp4 and gif), APK sideloading, bug reports, and custom scrip…
235414maintenance
CUT3R/CUT3R
CUT3R is the official PyTorch implementation of 'Continuous 3D Perception Model with Persistent State' (CVPR 2025 Oral), a stateful recurre…
381489active
pur1fying/blue_archive_auto_script
BAAS (Blue Archive Auto Script) is a GUI-based automation program for the mobile game Blue Archive that runs against 16:9 emulator screens.…
721483active
Nutlope/napkins
Napkins.dev is an open-source web application that turns screenshots or wireframes of website designs into working React + Tailwind code us…
631476active
mit-han-lab/torchsparse
TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a…
261472active
sczhou/Upscale-A-Video
Upscale-A-Video is a diffusion-based model for real-world video super-resolution that takes low-resolution videos and text prompts as input…
261471active
autonomousvision/mip-splatting
Mip-Splatting is a research implementation of alias-free 3D Gaussian Splatting, introducing a 3D smoothing filter and 2D Mip filter to elim…
271468active
Habrador/Computational-geometry
A C# computational geometry library for Unity implementing intersection algorithms, triangulations (Delaunay, constrained Delaunay), Vorono…
321467active
laochiangx/Common.Utility
A large collection of C# utility and helper classes covering common .NET development tasks such as Excel/CSV/PDF handling, HTTP requests, J…
325306maintenance
bludit/bludit
Bludit is a free, open-source flat-file CMS written in PHP for building websites and blogs in seconds. It stores content as JSON files, req…
941465active
nk-o/jarallax
Jarallax is a dependency-free JavaScript library that adds smooth parallax scrolling effects to background images, <img> tags, inline eleme…
761459active
valentinfrlch/ha-llmvision
LLM Vision is a Home Assistant integration (installed via HACS) that uses multimodal large language models to analyze images, videos, live …
901457active
NVlabs/Fast-FoundationStereo
Fast-FoundationStereo is NVIDIA's official PyTorch implementation of a real-time zero-shot stereo matching model family, accepted to CVPR 2…
541454active
dbolya/yolact
YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++…
515241maintenance
zylo117/Yet-Another-EfficientDet-Pytorch
A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim…
235238maintenance
tianweiy/DMD2
DMD2 is the official PyTorch implementation of Improved Distribution Matching Distillation, a NeurIPS 2024 method that distills diffusion m…
281448active
Topdu/OpenOCR
OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta…
581440active
v-modal/vmodal_sdk_flutter
VModal for Flutter is a Dart/Flutter SDK that adds multimodal video search (semantic, ASR, and OCR based) and streamed video uploads to And…
571439active
jrzaurin/pytorch-widedeep
A PyTorch library for multimodal deep learning that combines tabular data with text and images using Wide and Deep model architectures. It …
621416active
Arthi-chaud/Meelo
Meelo is a self-hosted music streaming server designed for music collectors, similar to Plex or Jellyfin but focused on music. It offers ri…
991406active
Zejun-Yang/AniPortrait
AniPortrait is a Python framework from Tencent that generates photorealistic portrait animations from an audio clip and a reference image, …
255021maintenance
Nativ
Nativ is a free, MIT-licensed macOS desktop application for running open AI models locally on Apple Silicon Macs, built on MLX-VLM. It prov…
801397active
zju3dv/street_gaussians
Street Gaussians is a research implementation of the ECCV 2024 paper 'Modeling Dynamic Urban Scenes with Gaussian Splatting', which reconst…
401391active
Junyi42/monst3r
MonST3R is the official PyTorch implementation of an ICLR 2025 paper that estimates per-timestep geometry (pointmaps) from dynamic videos i…
361387active
arkohut/pensieve
Pensieve is a privacy-focused passive screen recording application that automatically takes screenshots of your screens, indexes them with …
851386active
gettalong/hexapdf
HexaPDF is a pure Ruby library with an accompanying CLI application for creating, manipulating, merging, encrypting, signing, and optimizin…
761382active
aws-samples/generative-ai-use-cases
Generative AI Use Cases (GenU) is a well-architected sample application from AWS that demonstrates business use cases for generative AI, bu…
921381active
yanx27/Pointnet_Pointnet2_pytorch
A pure PyTorch implementation of the PointNet and PointNet++ deep learning architectures for point cloud processing. It includes training a…
324945maintenance
autonomousvision/unimatch
UniMatch is a PyTorch research library implementing a unified transformer-based model for optical flow, stereo matching, and depth estimati…
321379stable
Meituan-AutoML/MobileVLM
MobileVLM is a family of compact vision language models (1.4B-3B parameters) designed to run efficiently on mobile devices, combining small…
171370active
PurpleDoubleD/locally-uncensored
Locally Uncensored is a free, open-source desktop AI studio (built with Tauri/TypeScript) that bundles uncensored local chat, a coding agen…
811366active
79E/ChatGpt-Web
A commercially-viable ChatGPT web application built with React and TypeScript, featuring a full admin backend for managing users, tokens, p…
261363active
Sense-X/Co-DETR
Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg…
321360stable
ant-research/CoDeF
CoDeF is the official PyTorch implementation of Content Deformation Fields, a video representation combining a canonical content field and …
284846maintenance
pthom/imgui_bundle
Dear ImGui Bundle is a batteries-included immediate-mode GUI framework for both Python and C++, built on Dear ImGui with integrated plottin…
941353active
campsite/campsite
Campsite is an open-source team communication app combining posts, chat, video calls, and AI-generated summaries for distributed teams. Thi…
314827maintenance
huridocs/pdf-document-layout-analysis
A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl…
821349active
ParisNeo/lollms-webui
LoLLMs WebUI is a local, single-user web interface for running large language models and multimodal AI systems, supporting hundreds of mode…
714785maintenance
IrisRainbowNeko/genshin_auto_fish
A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen…
234759maintenance
LLaVA-VL/LLaVA-NeXT
LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision …
644716maintenance
open-edge-platform/geti
Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru…
981325active
yukkcat/gemini-business2api
A self-hosted gateway service that exposes Gemini Business through an OpenAI-compatible API, with multi-account load balancing and an admin…
581324active
Capsize-Games/airunner
AI Runner is a privacy-focused desktop application for running local AI models offline, combining an AI chat companion with voice conversat…
791314active
albermax/innvestigate
iNNvestigate is a Python toolbox providing a common interface and out-of-the-box implementations of many neural network explanation methods…
301308active
davidhampgonsalves/Life-Dashboard
A low-power personal dashboard that repurposes a jailbroken Kindle's e-ink screen to display daily information like weather, calendar, and …
501304active
wpilibsuite/allwpilib
WPILib is the official software library suite for programming robots in the FIRST Robotics Competition (FRC) and FIRST Tech Challenge (FTC)…
841301stable
dabit3/react-native-ai
React Native AI is a full-stack framework for building cross-platform mobile AI apps with React Native and an Express server proxy. It prov…
691297active
reg-viz/reg-suit
reg-suit is a command-line tool for visual regression testing that compares current screenshots against previous snapshots and generates HT…
791290active
TencentQQGYLab/ELLA
ELLA is an Efficient Large Language Model Adapter that equips text-to-image diffusion models with LLM-based text understanding via a Timest…
251290active
RoyalVane/CLAN
Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation…
321289stable
jordanrendric/claude-video-vision
A Claude Code plugin with an MCP server that gives Claude the ability to watch and understand videos by extracting frames via ffmpeg and tr…
581284active
AaronJackson/vrn
Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses…
324516maintenance
flutter-ml/google_ml_kit_flutter
A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa…
761278active
nianticlabs/monodepth2
Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra…
324500maintenance
leoxiaobin/deep-high-resolution-net.pytorch
Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai…
324480maintenance
DimitarPetrov/stegify
stegify is a Go command line tool and library for LSB (Least Significant Bit) steganography that hides any file inside images such as PNG a…
231266stable
Yuliang-Liu/MonkeyOCRv2
MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2…
581261active
MemeMeow-Studio/MemeMeow
MemeMeow is a self-hosted meme/sticker management and retrieval application that lets users find images by describing the desired scene in …
651260active
hanxi/cups-web
A web-based print management application built on CUPS that turns a home USB printer into an always-available network print service. It run…
831256active

← prev page 39 / 43 next →