Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
ogkalu2/comic-translate
An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language…
902911active
bytedeco/javacpp-presets
JavaCPP Presets provides Java bindings for commonly used native C++ libraries such as OpenCV, FFmpeg, and many others, built on the JavaCPP…
862850active
nut-tree/nut.js
nut.js is a cross-platform native UI automation and testing library for Node.js/TypeScript that controls mouse, keyboard, screen, and windo…
232847active
QwenLM/Qwen-MM-Plugins
A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.…
572777active
apple/turicreate
Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj…
1011159maintenance
AmberSahdev/Open-Interface
Open Interface is a cross-platform desktop application that lets users control their computer using natural language requests processed by …
652715active
Project-N-E-K-O/N.E.K.O
Project N.E.K.O. is an open-source AI companion application — a proactive catgirl-style AI that lives on your desktop, initiates interactio…
852681active
MrGiovanni/UNetPlusPlus
Official implementation of UNet++, a nested U-Net architecture for medical image segmentation, in both Keras and PyTorch. It redesigns skip…
772679stable
phillipi/pix2pix
The original Torch (Lua) implementation of pix2pix, a conditional GAN for image-to-image translation tasks such as synthesizing photos from…
3210652maintenance
heshengtao/super-agent-party
Super Agent Party is a self-hosted, all-in-one AI desktop companion that combines VRM-based virtual characters, agent skills, MCP tool supp…
872607active
VITA-MLLM/VITA
VITA is an open-source interactive omni multimodal large language model (VITA-1.5) that supports real-time vision and speech interaction, s…
292534active
facebookresearch/pifuhd
PIFuHD is a PyTorch implementation of a CVPR 2020 research model that reconstructs high-resolution 3D human body meshes from a single 2D im…
109737maintenance
wolny/pytorch-3dunet
A PyTorch implementation of 3D U-Net and its variants (residual, squeeze-and-excitation) for volumetric semantic segmentation, with 2D U-Ne…
632416active
HuCaoFighting/Swin-Unet
Official PyTorch implementation of Swin-Unet, a U-shaped pure Transformer model for medical image segmentation, published at ECCV 2022 Medi…
412416stable
ossappscollective/OSS-DocumentScanner
OSS Document Scanner is a free, open-source, privacy-focused mobile app for scanning documents with automatic edge detection, editing, OCR,…
912385active
ailia-ai/ailia-models
A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,…
772385active
zai-org/GLM-V
GLM-V is the open-source repository for Zhipu AI's GLM-4.6V, GLM-4.5V, and GLM-4.1V-Thinking vision-language models, which perform versatil…
602370active
Turbo1123/roubao
Roubao is an open-source AI phone automation assistant for Android, built natively in Kotlin and powered by vision-language models. It runs…
532325active
PKU-YuanGroup/MoE-LLaVA
MoE-LLaVA is an open-source Mixture-of-Experts based sparse large vision-language model, released with the MoE-Tuning training strategy fro…
312322active
apple/ml-ferret
Apple's Ferret, an end-to-end multimodal large language model (MLLM) that accepts any-form referring and grounds anything in its responses,…
278674maintenance
DanOps-1/Gpt-Agreement-Payment
A Python toolkit that reverse-engineers and replays the end-to-end ChatGPT Plus/Team/Pro subscription payment flow (Stripe Checkout, PayPal…
532225active
NVlabs/MambaVision
MambaVision is NVIDIA's official PyTorch implementation of a hybrid Mamba-Transformer vision backbone, published at CVPR 2025. It provides …
492224active
ellisdg/3DUnetCNN
A PyTorch library for building, training, and applying 3D U-Net convolutional neural networks for medical image segmentation. It provides c…
452224active
learnhouse/learnhouse
LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com…
992203active
kijai/ComfyUI-LivePortraitKJ
ComfyUI custom nodes that integrate the LivePortrait face animation and retargeting model, supporting image-to-video, video-to-video, and n…
232200active
cameroncooke/AXe
AXe is a Swift-based CLI tool for automating and inspecting iOS Simulators on macOS using Apple's private Accessibility APIs and HID input.…
812144active
PKU-YuanGroup/LLaVA-CoT
LLaVA-CoT is an 11B vision language model with training, inference, and dataset-generation code for spontaneous step-by-step multimodal rea…
472132active
vitoplantamura/OnnxStream
A lightweight C++ inference library for ONNX models that streams weights to run large models in very little memory, accelerated by XNNPACK.…
592086active
SizheAn/PanoHead
PanoHead is the official PyTorch implementation of a CVPR 2023 paper presenting a 3D-aware GAN that synthesizes geometry-aware, view-consis…
291956active
alibaba/EasyCV
EasyCV is an all-in-one PyTorch-based computer vision toolkit from Alibaba covering self-supervised learning, vision transformers, and majo…
321954active
microsoft/Magma
Magma is Microsoft Research's foundation model for multimodal AI agents, released as an 8B vision-language model that understands images an…
531937active
OpenTalker/video-retalking
VideoReTalking is a Python research system from SIGGRAPH Asia 2022 that edits real-world talking-head videos to match a given audio track, …
237280maintenance
IPADS-SAI/MobiAgent
MobiAgent is a systematic framework for building customizable GUI agents that operate mobile phones, comprising the MobiMind agent model fa…
601880active
NVIDIA-AI-IOT/Lidar_AI_Solution
NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil…
721867active
sarperavci/GoogleRecaptchaBypass
A Python library that automatically solves Google reCAPTCHA v2 challenges in under five seconds using browser automation with DrissionPage …
681855active
KlingAIResearch/ReCamMaster
ReCamMaster is a reference implementation of a camera-controlled generative video rendering model that re-renders a single source video alo…
441855active
NVIDIA/pix2pixHD
PyTorch implementation of pix2pixHD, a conditional GAN method for synthesizing and manipulating high-resolution (2048x1024) photorealistic …
326923maintenance
clovaai/donut
Donut is an OCR-free end-to-end Transformer model for visual document understanding tasks such as document classification and information e…
236919maintenance
GCWing/BitFun
BitFun is a cross-platform desktop AI agent application with a high-performance Rust agent runtime that writes code, produces documents, an…
821817active
yerfor/GeneFacePlusPlus
GeneFace++ is the official PyTorch implementation of a NeRF-based system for generalized and stable real-time 3D talking face generation. I…
261809active
zai-org/CogVLM
CogVLM is an open-source visual language model (17B) combining a vision encoder with a pretrained language model for image understanding an…
286744maintenance
OS-Copilot/OS-Copilot
OS-Copilot is an open-source Python library for building generalist AI agents that interface with operating system elements like the web, t…
161796active
QwenLM/Qwen-VL
Official repository for Qwen-VL, Alibaba Cloud's large vision-language model family, including the pretrained Qwen-VL and instruction-tuned…
286726maintenance
Webreaper/Damselfly
Damselfly is a server-based photograph management application designed to index and search very large image collections using metadata such…
881783active
xyTom/snippai
Snippai is an AI-powered snipping tool that captures screenshots and uses AI to extract structured content such as LaTeX formulas, text, ta…
891782active
google/automl
Google Brain's AutoML repository containing implementations of AutoML models and libraries such as EfficientNet, EfficientNetV2, and Effici…
106474maintenance
mahmoodlab/CLAM
CLAM is an open-source Python toolkit for data-efficient, weakly supervised classification of whole-slide images (WSIs) in computational pa…
391728active
facebookresearch/ConvNeXt
Official PyTorch implementation of ConvNeXt, a pure convolutional neural network architecture from the CVPR 2022 paper 'A ConvNet for the 2…
106416maintenance
AutoArk/EVA-OS
EVA OS / EVA Platform is a real-time multimodal AI operating system and development platform for next-generation smart hardware, combining …
681712active
robin-shaun/XTDrone
XTDrone is a customizable UAV simulation platform built on PX4, ROS, and Gazebo, supporting multi-rotors, fixed-wing, VTOL vehicles, and ot…
481711active
PaddlePaddle/PaddleVideo
PaddleVideo is a video understanding toolkit built on PaddlePaddle, offering state-of-the-art models for action recognition, temporal actio…
261702active
aisingapore/TagUI
TagUI is a free, open-source robotic process automation (RPA) tool from AI Singapore that lets users write simple text flows to automate re…
656325maintenance
mhamilton723/FeatUp
FeatUp is a model-agnostic framework that upsamples the spatial resolution of deep neural network features by 16-32x without changing their…
161654active
taki0112/UGATIT
Official TensorFlow implementation of U-GAT-IT, an unsupervised image-to-image translation model using attention modules and adaptive layer…
326116maintenance
ghostwright/ghost-os
Ghost OS is a native macOS framework that gives AI agents full computer-use capabilities by exposing the macOS accessibility tree, Chrome D…
611646active
CoinCheung/BiSeNet
A PyTorch implementation of the BiSeNet V1 and V2 real-time semantic segmentation models, with pretrained weights for Cityscapes, COCO-Stuf…
571638active
yoshitomo-matsubara/torchdistill
torchdistill is a modular, configuration-driven PyTorch framework for knowledge distillation and general deep learning experiments, requiri…
861629active
ZiqiaoPeng/SyncTalk
SyncTalk is the official PyTorch implementation of a CVPR 2024 paper that synthesizes speech-driven, synchronized talking head videos using…
461626active
open-gigaai/giga-world-0
GigaWorld-0 is a unified world model framework that acts as a data engine for Vision-Language-Action (VLA) learning in embodied AI. It comb…
411612active
dmlc/gluon-cv
GluonCV is a deep learning toolkit providing state-of-the-art computer vision model implementations with 170+ pre-trained models. It suppor…
235916maintenance
pq-yang/MatAnyone
MatAnyone is a CVPR 2025 human video matting framework that extracts alpha mattes of target people from video using consistent memory propa…
541605active
ml4a/ml4a
ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li…
321602active
Roy3838/Observer
Observer AI is a desktop application for building micro-agents that observe screen, camera, microphone, and audio inputs, process them with…
871601active
semperai/amica
Amica is an open-source web application for conversing with customizable 3D characters through voice chat, speech recognition, and vision. …
321593active
Tencent-Hunyuan/HunyuanWorld-Voyager
HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ…
521590active
Drexubery/ViewCrafter
ViewCrafter is a research codebase that uses video diffusion models to synthesize high-fidelity novel views of scenes from a single or spar…
491587active
BloodAxe/pytorch-toolbelt
A Python library of PyTorch extensions providing building blocks for fast R&D prototyping, including encoder-decoder architectures, special…
441574active
byjlw/video-analyzer
A Python CLI tool that analyzes videos by extracting key frames, transcribing audio with Whisper, and describing content using vision LLMs …
601559active
kritiksoman/GIMP-ML
GIMP-ML is a set of Python plugins that bring computer vision and deep learning models into the GNU Image Manipulation Program (GIMP). It p…
231553active
tianrun-chen/SAM-Adapter-PyTorch
A PyTorch library that adapts Meta AI's Segment Anything Model (SAM, SAM2, SAM3) to underperforming downstream segmentation tasks using lig…
671551active
cchen156/Learning-to-See-in-the-Dark
TensorFlow implementation of 'Learning to See in the Dark' (CVPR 2018), a deep learning model that brightens very dark, short-exposure RAW …
515565maintenance
WenmuZhou/PytorchOCR
A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP…
591523active
FeiYull/TensorRT-Alpha
A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl…
321460active
siddsachar/row-bot
Row-Bot is a local-first desktop AI assistant and workbench that combines chat, durable memory, a personal knowledge graph, tool use, paren…
811455active
ByteDance-Seed/m3-agent
M3-Agent is a multimodal agent framework from ByteDance Seed that processes real-time visual and auditory inputs to build entity-centric lo…
481445active
XPandora/PhysGaussian
PhysGaussian is a research library that integrates Material Point Method (MPM) physics simulation with 3D Gaussian Splatting representation…
551414active
nv-tlabs/GEN3C
GEN3C is NVIDIA's research codebase for a generative video model that achieves precise camera control and temporal 3D consistency using a 3…
591409active
sMythicalBird/ZenlessZoneZero-Auto
A Python-based automation framework for the game Zenless Zone Zero that uses image classification, template matching, and OCR to perform au…
221378active
anyrtcIO-Community/anyRTC-RTMP-OpenSource
anyLive is an open-source cross-platform live streaming SDK from anyRTC built on a WebRTC-93 base, providing RTMP push (publishing) and RTM…
304909maintenance
QwenLM/Qwen3-VL-Embedding
Qwen3-VL-Embedding and Qwen3-VL-Reranker are state-of-the-art multimodal embedding and reranking models built on the Qwen3-VL foundation mo…
561369active
ImprintLab/MedSegDiff
MedSegDiff is a diffusion probabilistic model framework for segmenting and reconstructing organs and tissues from medical images, with a tr…
501363active
a-real-ai/pywinassistant
PyWinAssistant is an open-source agentic framework that acts as a Computer-Using-Agent, operating Windows 10/11 graphical user interfaces e…
191343active
morettt/my-neuro
An open-source AI desktop companion framework inspired by Neuro-sama, letting users build a customizable Live2D character with sub-second v…
841342active
ZJU-REAL/ClawGUI
ClawGUI is a unified Python framework for GUI agents covering the full lifecycle: online reinforcement learning training (ClawGUI-RL with G…
701338active
GauravSingh9356/J.A.R.V.I.S
A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,…
481332active
ImprintLab/Medical-SAM-Adapter
Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei…
391322active
huawei-noah/Efficient-Computing
A collection of efficient deep learning methods from Huawei Noah's Ark Lab, covering model compression, knowledge distillation, pruning, qu…
321307active
robodhruv/visualnav-transformer
Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie…
191294active
jbarrow/commonforms
CommonForms is a Python package and CLI that uses trained object-detection models (FFDNet-S/L) to automatically detect form fields in a PDF…
651284active
dcharatan/pixelsplat
pixelSplat is a PyTorch implementation of a feed-forward model that reconstructs 3D radiance fields parameterized by 3D Gaussian primitives…
271274stable
Visual-Agent/DeepEyes
DeepEyes is a research project that trains multimodal vision-language models to 'think with images' using end-to-end reinforcement learning…
441271active
amaiya/ktrain
ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul…
251268active
nv-tlabs/Difix3D
Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau…
321266active
stardist/stardist
StarDist is a Python library for object detection and instance segmentation in 2D and 3D microscopy images using star-convex shapes, built …
601255stable
metadriverse/metadrive
MetaDrive is an open-source, lightweight driving simulator built for AI and autonomy research, supporting compositional scene synthesis and…
391235active
facebookresearch/home-robot
HomeRobot is an open-source robotics stack from Meta AI for mobile manipulation tasks on low-cost hardware like the Hello Robot Stretch. It…
231234active
MoonshotAI/Kimi-VL
Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo…
331224active
mrousavy/react-native-fast-tflite
A high-performance TensorFlow Lite library for React Native built on Nitro Modules, using the low-level C/C++ TFLite core API with zero-cop…
841222active
frotms/PaddleOCR2Pytorch
A PyTorch port of PaddleOCR that lets you run PaddleOCR-trained models (detection, recognition, and document structure parsing) without the…
731205active
deepseek-ai/DeepSeek-VL
DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf…
254175maintenance

← prev page 14 / 16 next →