Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
bytedeco/javacpp-presets
JavaCPP Presets provides Java bindings for commonly used native C++ libraries such as OpenCV, FFmpeg, and many others, built on the JavaCPP…
862850active
OpenGVLab/InternImage
InternImage is a large-scale CNN-based vision foundation model that uses deformable convolutions as its core operator, released with pretra…
282841stable
openalpr/openalpr
OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete…
2311452maintenance
google-research/kubric
Kubric is a data generation pipeline from Google Research for creating semi-realistic synthetic multi-object videos with rich annotations l…
602808active
microsoft/MoGe
MoGe is a deep learning model from Microsoft Research that recovers 3D geometry from a single open-domain image, predicting metric point ma…
662807active
Open-Cascade-SAS/OCCT
Open CASCADE Technology (OCCT) is an open-source C++ development platform for 3D surface and solid modeling, CAD data exchange, and visuali…
942800stable
nutonomy/nuscenes-devkit
The official Python devkit for the nuScenes dataset, a large-scale autonomous driving dataset from Motional. It provides dataset loading, v…
662796stable
rom1504/clip-retrieval
A Python toolkit for computing CLIP embeddings for images and text and building a semantic search/retrieval system on top of them. It inclu…
572795active
imanoop7/Ollama-OCR
A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports…
262780active
QwenLM/Qwen-MM-Plugins
A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native.…
572777active
gem5/gem5
gem5 is the official open-source computer-system architecture simulator for research and teaching, modeling processor microarchitecture and…
882771active
autodistill/autodistill
Autodistill is a Python library that uses large foundation vision models (like Grounding DINO, Grounded SAM, and CLIP) to automatically lab…
292763active
NVlabs/stylegan2
The official TensorFlow implementation of StyleGAN2, NVIDIA's improved style-based generative adversarial network for high-quality uncondit…
3211184maintenance
NVIDIA/FastPhotoStyle
FastPhotoStyle is NVIDIA's official PyTorch implementation of the ECCV 2018 paper 'A Closed-form Solution to Photorealistic Image Stylizati…
2311177maintenance
kha-white/manga-ocr
Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to…
902758stable
apple/turicreate
Turi Create is a Python library from Apple that simplifies building custom machine learning models for tasks like image classification, obj…
1011159maintenance
jiangdongguo/AndroidUSBCamera
AUSBC (AndroidUSBCamera) is a flexible UVC (USB video class) camera engine for Android, refactored in Kotlin with native C components. It s…
232750active
infstellar/genshin_impact_assistant
A multi-functional Genshin Impact auto-assist application that uses image recognition and simulated keystrokes to automate combat, domain r…
102750active
viser-project/viser
Viser is a Python library for web-based 3D visualization aimed at computer vision and robotics. It provides APIs for rendering 3D primitive…
992747active
stevenlovegrove/Pangolin
Pangolin is a lightweight, portable C++ utility library for rapid prototyping of 3D, numeric, and video-based programs, providing cross-pla…
872739stable
CVCUDA/CV-CUDA
CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs…
932718active
hiukim/mind-ar-js
MindAR is a web augmented reality library supporting image tracking and face tracking, written end-to-end in JavaScript with TensorFlow.js.…
232718active
lengstrom/fast-style-transfer
A TensorFlow implementation of fast neural style transfer that applies the style of famous paintings to photos and videos in real time. It …
3210962maintenance
bmild/nerf
The official TensorFlow implementation of NeRF (Neural Radiance Fields), the ECCV 2020 paper representing scenes as neural radiance fields …
3910927maintenance
teslamotors/react-native-camera-kit
A high-performance React Native camera library providing cross-platform camera capture, QR/barcode scanning, and face detection for iOS and…
962705active
yuweihao/MambaOut
MambaOut is a PyTorch implementation of Gated CNN models from the CVPR 2025 paper 'MambaOut: Do We Really Need Mamba for Vision?', which qu…
192704stable
magic-research/magic-animate
MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image …
4410897maintenance
TMElyralab/MusePose
MusePose is a diffusion-based, pose-guided image-to-video generation framework for creating virtual human videos, where a character in a re…
282701active
IDEA-Research/T-Rex
T-Rex is the official Python API client for T-Rex2, a generic open-set object detection model that combines text and visual prompts to dete…
482699active
dynobo/normcap
NormCap is an OCR-powered screen-capture application that lets users select a region of the screen and extracts its text to the clipboard i…
682695active
roboflow/maestro
maestro is a Python library from Roboflow that streamlines fine-tuning of multimodal vision-language models such as Florence-2, PaliGemma 2…
622694active
baaivision/EVA
EVA is a family of large-scale vision foundation models from BAAI, including masked image models (EVA-01/02) and scaled CLIP models (EVA-CL…
232691active
torinmb/mediapipe-touchdesigner
A GPU-accelerated, self-contained MediaPipe plugin for TouchDesigner that runs MediaPipe vision models (face detection, face/hand/pose trac…
862686active
bytedance/InfiniteYou
InfiniteYou (InfU) is a research framework from ByteDance for identity-preserved text-to-image generation built on Diffusion Transformers l…
372685active
MrGiovanni/UNetPlusPlus
Official implementation of UNet++, a nested U-Net architecture for medical image segmentation, in both Keras and PyTorch. It redesigns skip…
772679stable
HiLab-git/SSL4MIS
A benchmark and code collection of semi-supervised learning methods for medical image segmentation, re-implementing approaches like Mean Te…
442676active
tryolabs/norfair
Norfair is a lightweight, customizable Python library for real-time multi-object tracking that works with any detector outputting (x, y) co…
312676stable
JIA-Lab-research/LISA
LISA (Large Language Instructed Segmentation Assistant) is a multimodal large language model that performs reasoning-based image segmentati…
312674active
princeton-vl/DROID-SLAM
DROID-SLAM is a deep learning-based visual SLAM system for monocular, stereo, and RGB-D cameras that estimates camera trajectory and dense …
412671active
jlblancoc/nanoflann
nanoflann is a C++11 header-only library for fast nearest neighbor search using KD-trees. It is designed for efficient queries over point c…
942670stable
IceClear/StableSR
StableSR is a Python research library that leverages pre-trained Stable Diffusion priors for real-world blind image super-resolution. It pr…
212668stable
om-ai-lab/OmAgent
OmAgent is a Python library for building multimodal language agents, wrapping worker orchestration, task queues, and graph-based workflow o…
312665active
aigc3d/LHM
LHM is a PyTorch-based large reconstruction model that reconstructs high-fidelity animatable 3D human avatars from a single image in second…
522664active
phillipi/pix2pix
The original Torch (Lua) implementation of pix2pix, a conditional GAN for image-to-image translation tasks such as synthesizing photos from…
3210652maintenance
Tencent/MimicMotion
MimicMotion is a diffusion-based framework from Tencent for generating high-quality human motion videos guided by pose sequences, featuring…
472647active
colour-science/colour
Colour is an open-source Python library providing a comprehensive collection of colour science algorithms and datasets, including colour sp…
742642active
ultralytics/yolov3
Ultralytics' PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection. It provides training, validation…
6710596maintenance
anliyuan/Ultralight-Digital-Human
An ultralight talking-head (digital human) model that animates a person's face from audio input and runs in real time on mobile devices. It…
642627active
swz30/Restormer
Restormer is an efficient Transformer architecture for high-resolution image restoration, published as a CVPR 2022 Oral paper. It provides …
442625stable
OpenStitching/stitching
A Python package providing fast and robust image stitching to create panoramas, built on OpenCV's stitching module. It offers both a Python…
882620active
antoinelame/GazeTracking
A Python library that provides webcam-based eye tracking, returning pupil coordinates and gaze direction in real time using OpenCV and dlib…
682619active
luca-medeiros/lang-segment-anything
A Python library combining Meta's Segment Anything Model 2 with GroundingDINO to generate segmentation masks for objects in images specifie…
422598active
Slicer/Slicer
3D Slicer is a free, open-source desktop platform for visualization, processing, segmentation, registration, and analysis of medical and bi…
672595stable
Mininglamp-AI/Mano-P
Mano-P is an open-source GUI-VLA (vision-language-action) agent model and SDK for edge devices, enabling purely vision-driven cross-platfor…
552589active
kavan010/black_hole
A C++ black hole simulation that uses ray tracing and GPU compute shaders to render gravitational lensing, accretion disks, and spacetime c…
542588active
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0 is Tencent Hunyuan's open-source multi-modal world model framework that reconstructs, generates, and simulates 3D worlds from …
582571active
raulmur/ORB_SLAM2
ORB-SLAM2 is a real-time SLAM library for monocular, stereo, and RGB-D cameras that computes camera trajectories and sparse 3D reconstructi…
3210223maintenance
advimman/lama
LaMa is a PyTorch-based image inpainting model that fills large missing regions in images using fast Fourier convolutions, generalizing wel…
3410217maintenance
jolibrain/deepdetect
DeepDetect is an open-source deep learning runtime, CLI, and REST server written in C++ for training and inference across images, text, tab…
952551active
jpcy/xatlas
xatlas is a small C++11 library with no external dependencies that generates unique texture coordinates (UV unwrapping) for 3D meshes. It i…
322547stable
X-PLUG/mPLUG-Owl
mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and…
362539active
VITA-MLLM/VITA
VITA is an open-source interactive omni multimodal large language model (VITA-1.5) that supports real-time vision and speech interaction, s…
292534active
Intent-Lab/VisionClaw
VisionClaw is a real-time AI assistant app for Meta Ray-Ban smart glasses that streams camera frames and microphone audio to the Gemini Liv…
592529active
rpautrat/SuperPoint
A TensorFlow (with PyTorch conversion) implementation of the SuperPoint self-supervised interest point detector and descriptor network. It …
412511stable
luanfujun/deep-photo-styletransfer
Reference implementation of the CVPR 2017 paper 'Deep Photo Style Transfer', performing photorealistic image style transfer using Torch wit…
329989maintenance
BrunoLevy/geogram
Geogram is a C++ programming library of geometric algorithms for geometry processing, including surface reconstruction, remeshing, Boolean …
902499stable
yformer/EfficientSAM
EfficientSAM is an efficient image segmentation model that leverages masked image pretraining to provide a lightweight alternative to Meta'…
272491active
sthalles/SimCLR
A PyTorch reference implementation of SimCLR, a self-supervised contrastive learning framework for learning visual representations from unl…
232491stable
ppogg/YOLOv5-Lite
YOLOv5-Lite is a lightweight object detection model family evolved from YOLOv5, with models as small as ~900KB (int8) that run 10-15+ FPS o…
232487active
ipazc/mtcnn
A Python library implementing the MTCNN (Multitask Cascaded Convolutional Networks) algorithm for face detection and facial landmark alignm…
232485stable
twistedfall/opencv-rust
Rust bindings for the OpenCV computer vision library, generated automatically via Clang. It exposes OpenCV 3.4 (deprecated), 4.x, and 5.x A…
752483active
QIN2DIM/hcaptcha-challenger
A Python library that solves hCaptcha challenges using multimodal large language models and ONNX vision models (YOLO, CLIP, ResNet) integra…
862482active
xuebinqin/U-2-Net
Official PyTorch implementation of U^2-Net, a nested U-structure deep network for salient object detection, published in Pattern Recognitio…
329853maintenance
AprilRobotics/apriltag
AprilTag is a small C library implementing a visual fiducial (marker) detection system, computing the precise 3D position, orientation, and…
712480stable
GaParmar/img2img-turbo
A research library implementing one-step image-to-image translation models (CycleGAN-Turbo and pix2pix-turbo) built on SD-Turbo diffusion m…
412476active
wgsxm/PartCrafter
PartCrafter is a structured 3D generative model that jointly generates multiple semantically meaningful 3D mesh parts and objects from a si…
532471active
AngusJohnson/Clipper2
Clipper2 is a polygon clipping, offsetting, and triangulation library available in C++, C#, and Delphi. It performs boolean operations (int…
772462active
kevmo314/magic-copy
Magic Copy is a browser extension (Chrome, Firefox, and Figma) that uses Meta's Segment Anything Model to segment a foreground object from …
202458active
facebookresearch/pifuhd
PIFuHD is a PyTorch implementation of a CVPR 2020 research model that reconstructs high-resolution 3D human body meshes from a single 2D im…
109737maintenance
homuler/MediaPipeUnityPlugin
A Unity native plugin that ports the MediaPipe C++ API to C#, enabling MediaPipe graphs and solutions to run inside Unity applications. It …
632452active
hku-mars/r3live
R3LIVE is a tightly-coupled LiDAR-Inertial-Visual sensor fusion framework for robust, real-time state estimation and RGB-colored 3D mapping…
482442active
wangshub/Douyin-Bot
A Python bot that automates the Douyin (TikTok China) mobile app via ADB, taking screenshots and calling a face-recognition API to auto-lik…
329631maintenance
roboflow/inference
Roboflow Inference is a Python library and self-hostable inference server for deploying computer vision models on any computer or edge devi…
912427active
schappim/macOCR
macOCR is a macOS command-line tool that captures a screen region you select and runs OCR on it, copying the recognized text (or QR/barcode…
882426active
jyjblrd/Low-Cost-Mocap
A low-cost, room-scale motion capture system built with PlayStation cameras and ESP32 hardware, used to track objects and autonomously fly …
282421active
wolny/pytorch-3dunet
A PyTorch implementation of 3D U-Net and its variants (residual, squeeze-and-excitation) for volumetric semantic segmentation, with 2D U-Ne…
632416active
HuCaoFighting/Swin-Unet
Official PyTorch implementation of Swin-Unet, a U-shaped pure Transformer model for medical image segmentation, published at ECCV 2022 Medi…
412416stable
Smorodov/Multitarget-tracker
A C++ library for multiple object tracking that combines detectors (YOLO, D-FINE, RF-DETR, MobileNet-SSD) with tracking algorithms based on…
692415active
X-PLUG/mPLUG-DocOwl
mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO…
392411active
nv-tlabs/3dgrut
NVIDIA's official implementations of 3D Gaussian Ray Tracing (3DGRT) and 3D Gaussian Unscented Transform (3DGUT), which render volumetric G…
712390active
Alibaba-Quark/LiveAvatar
LiveAvatar is an open-source implementation of an ECCV 2026 paper for streaming, real-time, infinite-length audio-driven avatar video gener…
612386active
ailia-ai/ailia-models
A collection of 400+ pre-trained state-of-the-art AI models (object detection, pose estimation, speech recognition, LLMs, image generation,…
772385active
Cicada000/VV
A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d…
352375active
pixpark/gpupixel
GPUPixel is a high-performance, cross-platform real-time image and video filter library written in C++11 and built on OpenGL/ES. It provide…
882371active
zai-org/GLM-V
GLM-V is the open-source repository for Zhipu AI's GLM-4.6V, GLM-4.5V, and GLM-4.1V-Thinking vision-language models, which perform versatil…
602370active
OpenGVLab/InternVideo
InternVideo is a series of open-source video foundation models for multimodal video understanding, spanning generative and discriminative l…
722368active
tencent-ailab/V-Express
V-Express is a Python research project from Tencent AI Lab that generates talking head portrait videos from a reference image, audio, and V…
252360active
facebookresearch/perception_models
Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan…
542353active
MIT-SPARK/TEASER-plusplus
TEASER++ is a fast and certifiably-robust C++ library for rigid body point cloud registration in 3D, with Python and MATLAB bindings. It es…
482337stable
MouseLand/cellpose
Cellpose is a generalist deep learning algorithm for cellular and nucleus segmentation in microscopy images, with human-in-the-loop capabil…
862331active

← prev page 5 / 24 next →