Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: image-processing

4273 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
ingra14m/Deformable-3D-Gaussians
Official PyTorch implementation of the CVPR 2024 paper 'Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction'. …
181256stable
HisAtri/LrcApi
A Flask-based API service that fetches LRC lyrics and album/artist cover art for music players, built primarily for StreamMusic and Navidro…
701251active
Linketic/CityGaussian
Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia…
661251active
LaiFengiOS/LFLiveKit
LFLiveKit is an open-source RTMP live streaming SDK for iOS written in Objective-C. It provides H264/AAC hardware encoding, GPUImage beauty…
324391maintenance
ryokun6/ryos
ryOS is a web-based desktop environment that recreates classic macOS and Windows interfaces in the browser, built with React and TypeScript…
841244active
marian42/mesh_to_sdf
A Python library that computes approximate signed distance fields (SDFs) for arbitrary triangle meshes, including non-watertight, self-inte…
321242stable
showlab/Tune-A-Video
Tune-A-Video is the official PyTorch implementation of an ICCV 2023 paper that fine-tunes pre-trained text-to-image diffusion models (like …
314363maintenance
alibaba/Tora
Tora is Alibaba's official implementation of a trajectory-oriented Diffusion Transformer (DiT) for controllable video generation, integrati…
641241active
0xacx/chatGPT-shell-cli
A lightweight shell script that lets you chat with OpenAI's ChatGPT models and generate DALL-E images directly from the terminal, requiring…
321241active
totaljs/framework
Total.js framework is a full-featured, dependency-free web framework for Node.js written in pure JavaScript, comparable to Laravel, Django,…
234357maintenance
KonghaYao/cn-font-split
A Rust-based font subsetting tool that splits large CJK and other character fonts (otf, ttf, woff2) into small, web-ready packages with fin…
781239active
owen2345/camaleon-cms
Camaleon CMS is a dynamic and advanced content management system built on Ruby on Rails, designed as a flexible alternative to WordPress fo…
901238active
JamesHeinrich/getID3
getID3() is a PHP library that extracts metadata and technical information from a wide range of multimedia files, including audio, video, a…
821237stable
kellyvv/PhoneClaw
PhoneClaw is a mobile-native local AI agent framework that turns phones into on-device agent runtimes, running Gemma models via LiteRT and …
761232active
TheSmallHanCat/sora2api
A self-hosted OpenAI-compatible API gateway that wraps Sora's text-to-video and image generation capabilities behind standard /v1/chat/comp…
101232active
Tencent-Hunyuan/HunyuanCustom
HunyuanCustom is a multimodal-driven customized video generation framework built on HunyuanVideo, supporting image, text, audio, and video …
401228active
CharlyKeleb/SocialMedia-App
Wooble is a fully functional social media application built with Flutter and Dart, featuring photo feeds, real-time messaging, stories, and…
381221active
MotrixLab/SMPLer-X
Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an…
591219stable
fudan-generative-vision/champ
Champ is a research framework for controllable and consistent human image animation using 3D parametric guidance (SMPL-based depth, normal,…
254263maintenance
thClaws/thClaws
thClaws is an open-source AI agent harness written in native Rust that ships as a single binary offering a desktop GUI, CLI, headless, and …
771214active
ifzhang/FairMOT
FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification…
324245maintenance
Picsart-AI-Research/Text2Video-Zero
Official implementation of Text2Video-Zero, a zero-shot text-to-video generation method that adapts text-to-image diffusion models like Sta…
304243maintenance
LifeArchiveProject/BilibiliHistoryFetcher
A Python/FastAPI backend tool that fetches, stores, and analyzes a user's Bilibili watch history, favorites, dynamics, comments, and intera…
811205active
yanchunhuo/AutomationTest
A Python-based automation testing framework supporting API automation, web UI automation (Selenium), app UI automation (Appium), and perfor…
621205active
Artelnics/opennn
OpenNN is an open-source C++ library for building, training, and deploying neural networks for advanced analytics. It is dependency-free, o…
971199active
EvolvingLMMs-Lab/LLaVA-OneVision-2
A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis…
721197active
Tencent-Hunyuan/HunyuanWorld-Mirror
HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds…
541195active
zai-org/VisualGLM-6B
VisualGLM-6B is an open-source multimodal conversational language model supporting images, Chinese, and English, built on ChatGLM-6B with a…
304155maintenance
goodreasonai/ScrapeServ
ScrapeServ is a self-hosted API service that accepts a URL and returns the website's data along with browser screenshots, using Playwright …
241181active
Siv3D/OpenSiv3D
Siv3D (formerly OpenSiv3D) is a C++20 framework for creative coding, supporting 2D/3D games, media art, visualizers, and simulators. It pro…
671180active
DAMO-NLP-SG/VideoLLaMA3
VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de…
371178active
balancap/SSD-Tensorflow
A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net…
324101maintenance
DavidVentura/offline-translator
An Android app that translates text, PDF/ODT documents, and images entirely offline using Firefox translation models on-device. It also off…
881177active
Tencent-Hunyuan/MixGRPO
MixGRPO is a research framework from Tencent Hunyuan implementing a mixed ODE-SDE GRPO algorithm for efficient reinforcement learning fine-…
581177active
tjiiv-cprg/EPro-PnP
EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc…
401175stable
keinsaasforever/better-chatbot
Keinsaas Navigator (formerly Better Chatbot) is an open-source, self-hostable AI chatbot workspace built with Next.js and the Vercel AI SDK…
711172active
magicleap/SuperGluePretrainedNetwork
SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t…
324078maintenance
jenly1314/MLKit
MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin…
781171active
whyiyhw/chatgpt-wechat
A self-hosted Go application that lets users safely use LLM assistants (ChatGPT, Gemini, DeepSeek, Dify workflows) inside WeChat by relayin…
581170active
sergix44/xbackbone
XBackBone is a self-hosted, lightweight file and media sharing platform with first-class ShareX support. It provides a web UI, multi-user m…
831167active
zai-org/SCAIL-2
Official implementation of SCAIL-2, an open-source model for end-to-end controlled character animation that drives character videos from re…
581166active
sirius-ai/LPRNet_Pytorch
A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus…
321159stable
fundamentalvision/Deformable-DETR
Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl…
324018maintenance
cloudflare/ai
A monorepo of TypeScript packages and examples for building AI-powered applications on Cloudflare. It provides Vercel AI SDK and TanStack A…
791154active
OpenGVLab/VisionLLM
VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc…
331154active
GENEXIS-AI/chromex
Chromex is a Chrome MV3 side-panel extension that connects the browser to OpenAI's Codex CLI via a local native-messaging bridge. It lets u…
721153active
Woolverine94/biniou
biniou is a self-hosted web UI for 30+ generative AI models covering image, video, audio, and text generation, built with Gradio and Huggin…
631150active
amazon-science/mm-cot
Official PyTorch implementation of the paper 'Multimodal Chain-of-Thought Reasoning in Language Models', which adds vision features to a tw…
313985maintenance
magicrew/doc7
doc7 is a Go CLI tool that converts PDFs, Office files, scans, screenshots, charts, formulas, and diagrams into AI-ready Markdown using any…
771148active
cvg/glue-factory
Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point…
691143active
laravel/ai
The Laravel AI SDK is a PHP package offering a unified, expressive API for interacting with AI providers such as OpenAI, Anthropic, and Gem…
841140active
Anionex/agent-vision-toolkit
A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, …
791140active
HengyiWang/spann3r
Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.…
261140active
ifengzp/cocos-awesome
A collection of commonly used game feature modules and shader effect implementations for the Cocos Creator game engine, written in TypeScri…
351139active
clovaai/deep-text-recognition-benchmark
Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio…
323941maintenance
spacedeck/spacedeck-open
Spacedeck Open is a free, open-source, web-based collaborative whiteboard application with rich media support, originally a commercial SaaS…
321135active
weigert/TinyEngine
A small C++ OpenGL wrapper and 2D/3D rendering engine (~2000 lines of code) that abstracts boilerplate OpenGL for windows, shaders, texture…
231129active
geopavlakos/hamer
HaMeR (Hand Mesh Recovery) is a transformer-based model that reconstructs 3D hand meshes from single monocular images using the MANO parame…
561128active
zsyggg/paper-craft-skills
A collection of Claude Code / Codex agent skills that turn academic papers into method figures, visual slide decks, and in-depth HTML artic…
531128active
sml2h3/ddddocr-fastapi
A minimal FastAPI-based REST API service wrapping the DdddOcr OCR engine, exposing endpoints for image text recognition, slide captcha matc…
231126active
OpenGVLab/VideoMamba
VideoMamba is a state space model (Mamba-based) architecture for efficient video understanding, released with code and pretrained models fr…
251125active
FlagAI-Open/FlagAI
FlagAI is a Python toolkit for training, fine-tuning, and deploying large-scale AI models across NLP, CV, and vision-language tasks. It int…
643869maintenance
yangxue0827/RotationDetection
AlphaRotate is a TensorFlow-based benchmark and toolbox for rotated (oriented) object detection, implementing detectors such as R2CNN, Reti…
231118active
FutureUniant/Tailor
Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea…
371117active
HITsz-TMG/Uni-MoE
Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,…
681116active
THU-MIG/RepViT
Official PyTorch implementation of RepViT, a family of lightweight CNNs designed by integrating efficient ViT architectural designs into Mo…
191112stable
ATH-MaaS/Pixelle-MCP
Pixelle MCP is an open-source omnimodal AIGC framework that converts ComfyUI workflows (local or RunningHub cloud) into MCP tools with zero…
441109active
GPUOpen-LibrariesAndSDKs/Cauldron
Cauldron is a C++ framework by AMD for rapid prototyping of rendering demos and samples on Vulkan and Direct3D 12. It provides glTF 2.0 loa…
231098stable
farbrausch/fr_public
An archive of Farbrausch's demoscene tools from 2001-2011, including the Werkkzeug visual content-creation tools, the V2 synthesizer, kkrun…
323772maintenance
YOOTeam/OpenPPT
OpenPPT is a web-based online presentation (PPT) editor based on ChatPPT, supporting the full workflow of creating, importing, editing, bea…
361096active
Eyeline-Labs/Go-with-the-Flow
Official implementation of the CVPR 2025 Oral paper 'Go-with-the-Flow', which controls motion in video diffusion models by replacing i.i.d.…
421094active
yeahhe365/Gemini-Nexus
Gemini Nexus is a Chrome extension (Manifest V3) that adds an AI assistant layer to the browser, integrating Gemini Web, the Gemini API, an…
841092active
rhymes-ai/Aria
Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda…
231087active
7thSamurai/steganography
A C++ command-line tool that encrypts files with password-protected AES-256-CBC and hides them inside images using Least-Significant-Bit pi…
321085stable
1061700625/WeChat_Article
A PyQt5 desktop application that crawls and downloads all articles from a specified WeChat official account. It uses Selenium to log in and…
621078active
zju3dv/InfiniDepth
InfiniDepth is a CVPR 2026 research library for monocular depth estimation that represents depth as neural implicit fields, allowing depth …
521078active
symfony/ux
Symfony UX is an initiative and collection of PHP/JavaScript packages that integrate frontend tools like Stimulus, Turbo, Chart.js, React, …
981074active
orailnoor/cross-platform-llm-client
PrivateLM is a cross-platform AI chat client built with Flutter that unifies local on-device LLM inference (GGUF models with Vulkan GPU acc…
721074active
cleardusk/3DDFA
A PyTorch implementation of the TPAMI 2017 paper 'Face Alignment in Full Pose Range: A 3D Total Solution' (3DDFA). It fits a 3D Morphable M…
233677maintenance
microsoft/Biodiversity
Microsoft AI for Good Lab's biodiversity research hub providing open-source AI models and tools for wildlife monitoring and conservation, i…
881068active
gangweix/pixel-perfect-depth
Pixel-Perfect Depth is a monocular depth estimation model based on pixel-space diffusion transformers that produces flying-pixel-free depth…
491066active
open-gigaai/giga-models
GigaModels is an open-source Python framework providing pipelines for training, inference, deployment, and compression of multi-modal, gene…
611058active
YunYang1994/tensorflow-yolov3
A TensorFlow 1.x implementation of the YOLOv3 real-time object detector, reproducing the 'YOLOv3: An Incremental Improvement' paper. It sup…
233613maintenance
henry123-boy/SpaTracker
SpatialTracker is the official PyTorch implementation of a CVPR 2024 Highlight paper that tracks any 2D pixels in 3D space from RGB or RGBD…
411057active
ShenhanQian/GaussianAvatars
Official research code for GaussianAvatars, a CVPR 2024 Highlight method that creates photorealistic, fully controllable head avatars by ri…
561054active
showlab/MotionDirector
MotionDirector is a research library for customizing text-to-video diffusion models to generate videos with desired motions from a small se…
271053active
BAAI-DCAI/Bunny
Bunny is a family of lightweight multimodal vision-language models that combine plug-and-play vision encoders (EVA-CLIP, SigLIP) with langu…
261053active
buddhi1980/mandelbulber2
Mandelbulber v2 is a cross-platform desktop application for generating and rendering photorealistic three-dimensional fractals such as Mand…
711050active
agents-flex/agents-flex
Agents-Flex is a lightweight, modular Java framework for building AI applications and agents, positioned as a Java counterpart to Spring AI…
941046active
hanFengSan/eHunter
eHunter is a Tampermonkey/userscript that injects a Vue 3-based comic reader UI into supported comic sites (EH/EXHentai, NHentai), offering…
751045active
bs-community/blessing-skin-server
Blessing Skin is a self-hosted PHP web application for uploading, managing, and sharing custom Minecraft skins and capes, restoring them fo…
561044active
zju3dv/EfficientLoFTR
Efficient LoFTR is a PyTorch implementation of a semi-dense local feature matching model that matches keypoints between image pairs with sp…
401044active
EchoMimic
EchoMimic is a series of open-source models (V1-V3) from Ant Group for audio-driven human animation, generating lifelike talking-head, port…
501038active
liyupi/yu-picture
An enterprise-grade collaborative cloud image library platform built with Vue 3, Spring Boot, Tencent COS object storage, and WebSocket. It…
241038active
daniel-j/send2ereader
A self-hostable web service for sending ebooks (EPUB, MOBI, PDF, TXT, CBZ, CBR) to a Kobo or Kindle ereader via its built-in browser using …
391032active
jianjieyiban/JJYB_AI_VideoAutoCut
JJYB_AI 智剪 is a local-first desktop AI video creation workbench that combines material analysis, smart shot segmentation, commentary script…
701027active
vastxie/99AI
99AI is a commercially viable, self-hostable AI web platform built with Vue and Node.js that bundles AI chat, image/video/music generation,…
351027active
aim-uofa/AdelaiDet
AdelaiDet is an open-source Python toolbox built on Detectron2 that implements multiple instance-level detection and recognition algorithms…
323477maintenance
open-mmlab/mmyolo
MMYOLO is the OpenMMLab toolbox and benchmark for the YOLO series of object detection models, implemented on PyTorch. It provides unified i…
233468maintenance
MeshAnything
MeshAnything is an autoregressive transformer model that generates artist-created 3D meshes (up to 1600 faces in V2) aligned with a given s…
311018active

← prev page 40 / 43 next →