Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: image-processing

4273 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
pimcore/pimcore
Pimcore is an open core PHP framework and platform for Product Experience Management that unifies PIM, MDM, DAM, CDP, DXP/CMS, and digital …
953841active
google-research/scenic
Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr…
763821active
MaaEnd/MaaEnd
MaaEnd is a vision-AI-powered automation assistant for the game 'Arknights: Endfield', built on MaaFramework. It captures the screen, recog…
953762active
abhiTronix/vidgear
VidGear is a high-performance, cross-platform Python framework for video processing built around multi-threaded and asynchronous pipelines.…
763723active
facebookresearch/map-anything
MapAnything is an open-source research framework from Meta and CMU for universal feed-forward metric 3D reconstruction using an end-to-end …
773707active
ferdous-alam/GenCAD
GenCAD is a research codebase for image-conditioned CAD model generation using transformer-based contrastive representations (CCIP) and dif…
363676active
sukeesh/Jarvis
Jarvis is a command-line personal assistant for Linux, macOS, and Windows written in Python. It offers 15+ task categories including weathe…
483664active
Screenly/Anthias
Anthias (formerly Screenly OSE) is a free, open-source digital signage platform that turns a Raspberry Pi or x86 PC into a networked media …
993659active
MrNeRF/LichtFeld-Studio
LichtFeld Studio is a native open-source desktop application for 3D Gaussian Splatting that combines training, real-time inspection, splat …
923650active
EhViewer-NekoInverter/EhViewer
An Android client app for browsing E-Hentai/ExHentai galleries, forked from EhViewer with a classic Material Design 2 style. It is maintain…
843649active
NExT-GPT/NExT-GPT
NExT-GPT is an end-to-end any-to-any multimodal large language model that accepts and generates arbitrary combinations of text, image, vide…
373634active
GVCLab/PersonaLive
PersonaLive is a diffusion-based framework for real-time, streamable portrait image animation, generating infinite-length expressive talkin…
533627active
ZhaoJ9014/face.evoLVe
A high-performance face recognition library built on PaddlePaddle and PyTorch, providing comprehensive tools for face-related analytics and…
383591active
sligter/LandPPT
LandPPT is an AI-powered presentation generation platform that turns a topic or uploaded documents (PDF, Word, Markdown, Excel, PPT) into p…
813582active
Mereithhh/vanblog
VanBlog is a self-hosted personal blogging system built with Next.js/TypeScript that combines a static-site-generated frontend, an admin da…
343581active
liu-ziting/what-to-eat
An AI-powered recipe generation web platform built with Vue 3 and TypeScript that creates recipes across China's eight major cuisines plus …
453510active
mne-tools/mne-python
MNE-Python is an open-source Python library for exploring, visualizing, and analyzing human neurophysiological data such as MEG, EEG, sEEG,…
883502stable
PKU-YuanGroup/Video-LLaVA
Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into…
273500active
HeapsIO/heaps
Heaps is a high-performance, cross-platform 2D and 3D game engine and graphics framework written in Haxe, created by the designer of the Ha…
743499stable
huxingyi/dust3d
Dust3D is a free, open-source, cross-platform 3D modeling application for creating low-poly 3D models quickly. It automates UV unwrapping, …
953483active
roflcoopter/viseron
Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti…
983475active
HanaokaYuzu/Gemini-API
A reverse-engineered asynchronous Python client library for Google's Gemini web app (formerly Bard), published as gemini-webapi on PyPI. It…
923472active
huangjunsen0406/py-xiaozhi
py-xiaozhi is an open-source, cross-platform multimodal AI voice assistant client written in Python, compatible with the xiaozhi-esp32 ecos…
853463active
n3d1117/chatgpt-telegram-bot
A self-hosted Telegram bot written in Python that integrates with OpenAI's official ChatGPT, DALL·E, and Whisper APIs to answer questions, …
333463active
Grt1228/chatgpt-java
An unofficial Java SDK for the OpenAI API covering all official endpoints including chat completions (GPT-3.5/GPT-4), DALL-E image generati…
213420active
davidsandberg/facenet
A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC…
3214348maintenance
paper-design/shaders
Paper Shaders is a collection of zero-dependency HTML canvas/WebGL shader components for websites, available as vanilla JS and React npm pa…
673414active
diced/zipline
Zipline is a self-hosted file upload and URL shortening server with a feature-rich dashboard, designed as a ShareX-compatible upload target…
993396active
WongKinYiu/yolov7
Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques…
2314141maintenance
timerring/bilive
BILIVE is a Python application that records Bilibili live streams and danmaku 24/7, then automatically renders danmaku and AI-generated sub…
613280active
jixiaozhong/Sonic
Sonic is the official PyTorch implementation of the CVPR 2025 paper 'Sonic: Shifting Focus to Global Audio Perception in Portrait Animation…
493274active
deepdoctection/deepdoctection
deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c…
983257active
cocktailpeanut/fluxgym
FluxGym is a simple web UI for training FLUX LoRA models with low VRAM support (12GB/16GB/20GB). It combines the AI-Toolkit Gradio frontend…
653251active
imraywang/wewrite
WeWrite is a Python-based AI agent skill that automates the full WeChat Official Account content pipeline: topic selection from trending ne…
793230active
mit-han-lab/bevfusion
BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re…
103230stable
orhun/ratty
Ratty is a GPU-rendered terminal emulator written in Rust with Ratatui that supports inline 3D graphics alongside traditional 2D terminal r…
783219active
kerlomz/captcha_trainer
A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren…
553211active
jy0205/Pyramid-Flow
Pyramid Flow is the official PyTorch implementation of a training-efficient autoregressive video generation model based on pyramidal flow m…
223211active
Pointcept
Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan…
763206active
Bionus/imgbrd-grabber
Grabber is a highly customizable imageboard/booru browser and mass downloader that can fetch thousands of images from multiple booru source…
883192active
Rudrabha/Wav2Lip
Wav2Lip is the official research code for the ACM Multimedia 2020 paper 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In …
4513193maintenance
liyown/ai-trend-publish
TrendPublish is a TypeScript-based automated content pipeline for WeChat Official Accounts that scrapes multiple sources (Twitter/X, RSS, s…
823170active
Devolutions/IronRDP
IronRDP is a modular Rust implementation of the Microsoft Remote Desktop Protocol (RDP), providing PDU codecs, connection/session state mac…
973147active
darkzOGx/youtube-automation-agent
AgentTube is a self-hosted Node.js application that uses AI agents to run a YouTube channel end to end: researching topics, writing scripts…
823118active
theajack/cnchar
cnchar is a comprehensive TypeScript library for Chinese character processing, offering pinyin conversion, stroke counts, stroke order draw…
663087active
TanStack/ai
TanStack AI is a type-safe, provider-agnostic TypeScript SDK for building AI applications with streaming chat, tool calling, agents, struct…
803060active
sonos/tract
Tract is Sonos' tiny, self-contained neural-network inference engine written in Rust. It loads ONNX, TensorFlow/TFLite, and NNEF models, op…
993053active
off-grid-ai/OGAM
Off Grid AI (OGAM) is a cross-platform mobile and desktop application that runs AI entirely on-device: GGUF LLM chat with vision, Whisper s…
783045active
korlibs/korge
KorGE is a modern multiplatform game engine written entirely in Kotlin, built on top of the Korlibs multimedia stack. It targets JVM/Androi…
673042active
SharpAI/DeepCamera
DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r…
863038active
viperrcrypto/Siftly
Siftly is a self-hosted, local-first web application for organizing Twitter/X bookmarks into a searchable, categorized knowledge base. It r…
643028active
MeiGen-AI/MultiTalk
MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a …
562998active
Dicklesworthstone/llm_aided_ocr
A Python tool that converts scanned PDFs to text via Tesseract OCR, then uses LLMs (local or API-based like OpenAI/Anthropic) to correct OC…
712997active
leigest519/ScreenCoder
ScreenCoder is a UI-to-code generation system that converts screenshots or design mockups into clean, editable HTML/CSS using a modular mul…
622976active
marcotcr/lime
Lime (Local Interpretable Model-agnostic Explanations) is a Python library that explains the predictions of any machine learning classifier…
2312161maintenance
mahlernim/google-timeline-visualizer
An Android app (with an iPhone web app) that turns Google Maps Timeline (Location History) exports into an animated travel video. Users sel…
832957active
sunsmarterjie/yolov12
YOLOv12 is a PyTorch implementation of attention-centric real-time object detectors, published at NeurIPS 2025. It provides detection model…
592954active
iscyy/ultralyticsPro
A PyTorch-based collection of improved YOLO-family object detection models (YOLOv5 through YOLOv13, RT-DETR) with pluggable modules for bac…
482954active
Rust-SDL2/rust-sdl2
Rust bindings for the SDL2 multimedia library, wrapping low-level C APIs in idiomatic Rust. It provides access to graphics, audio, input, a…
652950active
rust-headless-chrome/rust-headless-chrome
A Rust library providing a high-level API to control headless Chrome or Chromium via the DevTools Protocol, serving as the Rust equivalent …
812948active
microsoft/table-transformer
Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un…
232943active
MacPaw/OpenAI
A community-maintained Swift package that wraps the OpenAI public API, supporting chat completions, responses, function calling, MCP tools,…
942940active
openrecall/openrecall
OpenRecall is an open-source, privacy-first digital memory tool that periodically takes screenshots of your screen, extracts text via local…
432937active
sherlockchou86/VideoPipe
VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates …
542934active
jeeliz/jeelizFaceFilter
A lightweight JavaScript/WebGL library for real-time face detection and tracking from a camera feed via WebRTC, designed for building augme…
462933active
boredazfcuk/docker-icloudpd
An Alpine Linux Docker container wrapping the iCloud Photos Downloader (icloudpd) utility for syncing iCloud photo libraries to a local ser…
702932active
zeas2/Kirikiroid2
Kirikiroid2 is a cross-platform port of the Kirikiri2/KirikiriZ visual novel game engine, allowing Kirikiri-based games to run on platforms…
232929active
InternLM/InternLM-XComposer
InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u…
382925active
JunChenMoCode/ChatGPT_JCM
A Vue2 + ElementUI web management interface that aggregates OpenAI API endpoints (models, chat, images, audio, fine-tuning, files) into a g…
222923active
Ovi/DummyJSON
DummyJSON is a free hosted fake REST API that serves placeholder JSON data (products, users, carts, posts, quotes, todos, recipes) for fron…
732911active
spliit-app/spliit
Spliit is a free and open-source web application for sharing and tracking group expenses, serving as an alternative to Splitwise. It is bui…
922907active
Saiyan-World/goku
Goku is a family of flow-based (rectified flow Transformer) foundation models for joint image and video generation, released by HKU and Byt…
232905active
TheSmallHanCat/flow2api
Flow2API is a self-hosted Python/FastAPI service that exposes an OpenAI- and Gemini-compatible API on top of Google Flow (VideoFX/ImageFX) …
602873active
nut-tree/nut.js
nut.js is a cross-platform native UI automation and testing library for Node.js/TypeScript that controls mouse, keyboard, screen, and windo…
232848active
OpenGVLab/InternImage
InternImage is a large-scale CNN-based vision foundation model that uses deformable convolutions as its core operator, released with pretra…
282846stable
Open-Cascade-SAS/OCCT
Open CASCADE Technology (OCCT) is an open-source C++ development platform for 3D surface and solid modeling, CAD data exchange, and visuali…
942831stable
youniaogu/MangaReader
A cross-platform manga reading app built with React Native for Android and iOS, with tablet support. It uses a plugin-based design to aggre…
512818active
QwenLM/Qwen-MM-Plugins
A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.…
572799active
rom1504/clip-retrieval
A Python toolkit for computing CLIP embeddings for images and text and building a semantic search/retrieval system on top of them. It inclu…
572795active
lucidrains/DALLE2-pytorch
A PyTorch implementation of OpenAI's DALL-E 2 text-to-image synthesis model, focusing on the diffusion prior network that predicts image em…
2311301maintenance
OpenDCAI/Paper2Any
Paper2Any is an open-source Python application that turns research papers, text, or topics into editable research figures, technical route …
562781active
FWGS/xash3d-fwgs
Xash3D FWGS is a cross-platform game engine forked from Xash3D, aimed at compatibility with the Half-Life (GoldSrc) engine while extending …
862757active
sepinf-inc/IPED
IPED is an open source digital forensic tool developed by the Brazilian Federal Police for processing and analyzing digital evidence from d…
832737active
torinmb/mediapipe-touchdesigner
A GPU-accelerated, self-contained MediaPipe plugin for TouchDesigner that runs MediaPipe vision models (face detection, face/hand/pose trac…
862715active
bcosca/fatfree
Fat-Free Framework (F3) is a lightweight PHP micro-framework condensed into a single file for building dynamic web applications quickly. It…
842715stable
YusufB5/ASCILINE
ASCILINE is a high-performance ASCII video rendering engine written in Python that converts video pixels into text-based representations. I…
582709active
TMElyralab/MusePose
MusePose is a diffusion-based, pose-guided image-to-video generation framework for creating virtual human videos, where a character in a re…
282703active
roboflow/maestro
maestro is a Python library from Roboflow that streamlines fine-tuning of multimodal vision-language models such as Florence-2, PaliGemma 2…
622695active
chrome-php/chrome
A PHP library for controlling headless Chrome/Chromium browsers via the DevTools protocol, supporting both synchronous and asynchronous usa…
892677active
JIA-Lab-research/LISA
LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati…
312673active
Tencent/MimicMotion
MimicMotion is a diffusion-based framework from Tencent for generating high-quality human motion videos guided by pose sequences, featuring…
472653active
UniversalMediaServer/UniversalMediaServer
Universal Media Server is a free, open-source DLNA, UPnP and HTTP(S) media server that streams or transcodes video, audio and images to TVs…
972645active
ultralytics/yolov3
Ultralytics' PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection. It provides training, validation…
6710602maintenance
morethanwords/tweb
Telegram Web K is the open-source TypeScript web client powering web.telegram.org/k/, based on the original Webogram and actively patched a…
772628active
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from …
582601active
UglyToad/PdfPig
PdfPig is a C#/.NET library for reading and extracting text, images, annotations, forms, and metadata from PDF files, ported from Apache PD…
912556active
jolibrain/deepdetect
DeepDetect is an open-source deep learning runtime, CLI, and REST server written in C++ for training and inference across images, text, tab…
952551active
X-PLUG/mPLUG-Owl
mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and…
362539active
BrunoLevy/geogram
Geogram is a C++ programming library of geometric algorithms for geometry processing, including surface reconstruction, remeshing, Boolean …
902527stable
kpcyrd/sn0int
sn0int is a semi-automatic OSINT framework and package manager written in Rust that enumerates attack surface by processing public informat…
602522active

← prev page 37 / 43 next →