Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Parskatt/RoMa
RoMa (romatch) is a Python library for robust dense feature matching between image pairs, estimating pixel-dense warps and reliable certain…
521293active
PantoMatrix/PantoMatrix
PantoMatrix is an open-source research project that generates 3D face and body animation from speech audio, including the EMAGE model for h…
331293active
streamlit/demo-self-driving
A Streamlit demo app that provides an interactive image browser for the Udacity self-driving-car dataset with realtime YOLO object detectio…
601290stable
RoyalVane/CLAN
Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation…
321289stable
hku-mars/livox_camera_calib
A C++/ROS tool from HKU MARS for automatic extrinsic calibration between high-resolution LiDAR (e.g., Livox) and cameras in targetless envi…
321289stable
huanngzh/MV-Adapter
MV-Adapter is a plug-and-play adapter that turns pre-trained text-to-image diffusion models (e.g., SDXL, SD2.1) into multi-view consistent …
341285active
Phantom-video/HuMo
HuMo is a research model and Python codebase from Tsinghua University and ByteDance for human-centric video generation using collaborative …
461283active
BoboTiG/python-mss
Python MSS is an ultra-fast, cross-platform screenshot library in pure Python using ctypes, capable of capturing one or all monitors with n…
811280stable
Tianxiaomo/pytorch-YOLOv4
A minimal PyTorch implementation of YOLOv4 (and YOLOv4-tiny) supporting inference and training, with tools to convert Darknet weights to Py…
324521maintenance
zju3dv/MatchAnything
MatchAnything is a deep learning model for universal cross-modality image matching, released as research code accompanying a TPAMI 2026 pap…
641279active
AaronJackson/vrn
Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses…
324517maintenance
plemeri/transparent-background
A Python tool and CLI that removes backgrounds from images and videos using the InSPyReNet deep learning model (ACCV 2022). It supports ima…
631278active
city-super/Scaffold-GS
Scaffold-GS is a research implementation of a structured 3D Gaussian splatting method that uses anchor points on a sparse voxel grid to dis…
271278active
nvidia-isaac/nvblox
nvblox is a GPU-accelerated C++/Python library for real-time 3D reconstruction using TSDF and ESDF volumetric mapping, designed for robots …
851276active
Fugtemypt123/VIGA
VIGA is an analysis-by-synthesis code agent that reconstructs 3D scenes and slide layouts from images by generating and executing Blender P…
551275active
Ma-Lab-Berkeley/CRATE
CRATE is the official PyTorch implementation of the Coding RAte reduction TransformEr, a family of 'white-box' transformer architectures de…
291275active
google-research/simclr
Google Research's official implementation of SimCLR and SimCLRv2, a framework for contrastive learning of visual representations, with 65 p…
104502maintenance
flutter-ml/google_ml_kit_flutter
A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa…
761274active
dcharatan/pixelsplat
pixelSplat is a PyTorch implementation of a feed-forward model that reconstructs 3D radiance fields parameterized by 3D Gaussian primitives…
271274stable
nianticlabs/monodepth2
Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra…
324497maintenance
NVlabs/stylegan2-ada-pytorch
Official PyTorch implementation of StyleGAN2-ADA, a generative adversarial network with adaptive discriminator augmentation for training wi…
324487maintenance
Visual-Agent/DeepEyes
DeepEyes is a research project that trains multimodal vision-language models to 'think with images' using end-to-end reinforcement learning…
441271active
leoxiaobin/deep-high-resolution-net.pytorch
Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai…
324480maintenance
AravisProject/aravis
Aravis is a C library based on GLib/GObject for video acquisition from Genicam-compliant industrial cameras, implementing GigE Vision and U…
811268active
BachiLi/diffvg
diffvg is a differentiable rasterizer for 2D vector graphics that bridges the raster and vector domains via backpropagation. It computes pi…
421268active
amaiya/ktrain
ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul…
251268active
ToniRV/NeRF-SLAM
NeRF-SLAM is a real-time dense monocular SLAM system that combines neural radiance fields (Instant-NGP) with probabilistic volumetric fusio…
321266active
nv-tlabs/Difix3D
Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau…
321266active
bryandlee/animegan2-pytorch
A PyTorch implementation of AnimeGANv2, a GAN-based image-to-image style transfer model that converts photos into anime-style images. It pr…
324452maintenance
ethz-asl/rovio
ROVIO (Robust Visual Inertial Odometry) is a C++ framework from ETH Zurich that estimates camera and IMU trajectory using an iterated exten…
611262active
rlawjdghek/StableVITON
StableVITON is the official PyTorch implementation of a CVPR 2024 paper that performs image-based virtual try-on using a pre-trained latent…
471261stable
Fictionarry/ER-NeRF
ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr…
241260stable
nv-tlabs/GET3D
GET3D is NVIDIA's PyTorch implementation of a generative model that synthesizes high-quality 3D textured meshes (cars, chairs, animals, bui…
324435maintenance
Yuliang-Liu/MonkeyOCRv2
MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2…
581256active
stardist/stardist
StarDist is a Python library for object detection and instance segmentation in 2D and 3D microscopy images using star-convex shapes, built …
601255stable
JonathonLuiten/TrackEval
TrackEval is a Python library for evaluating multi-object tracking (MOT) algorithms, implementing metrics such as HOTA, CLEARMOT, IDF1, VAC…
321255stable
ingra14m/Deformable-3D-Gaussians
Official PyTorch implementation of the CVPR 2024 paper 'Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction'. …
181255stable
roryclear/clearcam
Clearcam is a self-hosted Python NVR that adds AI object detection, tracking, mobile notifications, and semantic search to any RTSP securit…
861254active
ziyc/drivestudio
DriveStudio is a Python framework for 3D Gaussian Splatting (3DGS) based reconstruction and simulation of dynamic urban driving scenes. It …
401254active
VainF/pytorch-msssim
A PyTorch library providing fast, differentiable SSIM and MS-SSIM image quality metrics using separable Gaussian filtering for speed. It ca…
231253stable
FaceAISDK/FaceAISDK_Android
An Android SDK for fully on-device, offline face detection, recognition, liveness detection (anti-spoofing), and 1:1, 1:N, and M:N face sea…
981252active
Linketic/CityGaussian
Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia…
661251active
peterbraden/node-opencv
Native Node.js bindings for the OpenCV computer vision library, exposing Matrices, image reading/writing, and cascades like face detection …
234384maintenance
withoutbg/withoutbg-python
A Python SDK (pip install withoutbg) for removing image backgrounds, offering a free local open-weights ONNX model and an optional paid clo…
801246active
lpiccinelli-eth/UniDepth
UniDepth is a Python library and research codebase for universal monocular metric depth estimation from single images, based on CVPR 2024 a…
351246active
unum-cloud/UForm
UForm is a compact multimodal AI library providing tiny image-text embedding models (64-768 dimensions, Matryoshka-style) and small generat…
551244active
3DTopia/OpenLRM
OpenLRM is an open-source PyTorch implementation of Large Reconstruction Models (LRM) that reconstruct 3D objects (meshes and rendered vide…
171244active
abewley/sort
SORT is a barebones Python implementation of a simple online and realtime multiple object tracking algorithm for 2D video sequences, based …
324373maintenance
HJYao00/Mulberry
Mulberry is a research implementation of an o1-like multimodal large language model (MLLM) that performs step-by-step reasoning and reflect…
491243active
cvg/depthsplat
DepthSplat is a PyTorch research library implementing a CVPR 2025 model that connects Gaussian splatting with single/multi-view depth estim…
561242active
raspberrypi/picamera2
Picamera2 is a Python library providing an interface to Raspberry Pi cameras via the libcamera stack, replacing the legacy Picamera library…
941241active
alibaba/Tora
Tora is Alibaba's official implementation of a trajectory-oriented Diffusion Transformer (DiT) for controllable video generation, integrati…
641241active
XPixelGroup/HYPIR
Official PyTorch implementation of HYPIR, a SIGGRAPH 2025 method that harnesses diffusion-yielded score priors for image restoration. It pr…
391241active
jcjohnson/fast-neural-style
A Torch (Lua) implementation of feedforward neural style transfer from the ECCV 2016 paper 'Perceptual Losses for Real-Time Style Transfer …
324359maintenance
marian42/mesh_to_sdf
A Python library that computes approximate signed distance fields (SDFs) for arbitrary triangle meshes, including non-watertight, self-inte…
321240stable
ZHKKKe/MODNet
MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima…
324355maintenance
facebookresearch/deit
Official PyTorch repository for DeiT and related vision transformer architectures (CaiT, ResMLP, PatchConvnet, DeiT III), providing trainin…
104355maintenance
apchenstu/TensoRF
TensoRF is a PyTorch implementation of the ECCV 2022 paper 'TensoRF: Tensorial Radiance Fields', which models and reconstructs radiance fie…
441239stable
stella-cv/stella_vslam
stella_vslam is a community-maintained fork of OpenVSLAM implementing a monocular, stereo, and RGBD visual SLAM system in C++. It supports …
771238active
open-mmlab/playground
OpenMMLab Playground is a central hub collecting and showcasing community projects that extend OpenMMLab libraries with Segment Anything Mo…
301236active
bytedance/USO
USO is ByteDance's open-source unified style- and subject-driven image generation model based on diffusion (FLUX), combining any subject wi…
361235active
xtreme1-io/xtreme1
Xtreme1 is an open-source, self-hosted data labeling and annotation platform for multimodal training data, supporting images, 3D LiDAR poin…
621234active
facebookresearch/home-robot
HomeRobot is an open-source robotics stack from Meta AI for mobile manipulation tasks on low-cost hardware like the Hello Robot Stretch. It…
231234active
VAST-AI-Research/TripoSplat
TripoSplat is an inference-only Python library from TripoAI that converts a single 2D image into high-quality 3D Gaussian splats with a var…
571232active
stereolabs/zed-sdk
The ZED SDK is a cross-platform spatial perception library for Stereolabs ZED stereo cameras, providing depth sensing, SLAM, 3D reconstruct…
911229active
MoonshotAI/Kimi-VL
Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo…
331224active
mrousavy/react-native-fast-tflite
A high-performance TensorFlow Lite library for React Native built on Nitro Modules, using the low-level C/C++ TFLite core API with zero-cop…
841222active
maxritter/diy-thermocam
DIY-Thermocam is an open-source, self-assembly thermal imaging camera based on the FLIR Lepton sensor and a Teensy 4.1 microcontroller, wit…
481222active
fastgs/FastGS
FastGS is a general acceleration framework for 3D Gaussian Splatting that trains scenes in roughly 100 seconds using multi-view consistent …
491221active
MotrixLab/SMPLer-X
Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an…
591220stable
bowang-lab/MedRAX
MedRAX is a medical reasoning agent framework that integrates chest X-ray analysis tools (segmentation, grounding, report generation, disea…
421218active
fudan-generative-vision/champ
Champ is a research framework for controllable and consistent human image animation using 3D parametric guidance (SMPL-based depth, normal,…
254261maintenance
kijai/ComfyUI-segment-anything-2
A set of ComfyUI custom nodes that bring Meta's Segment Anything 2 (SAM2) models into ComfyUI workflows for promptable image and video segm…
431214active
Picsart-AI-Research/Text2Video-Zero
Official implementation of Text2Video-Zero, a zero-shot text-to-video generation method that adapts text-to-image diffusion models like Sta…
304245maintenance
ifzhang/FairMOT
FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification…
324244maintenance
DachunKai/EvTexture
Official PyTorch implementation of EvTexture and EvTexture++, event-driven video super-resolution models that use event-camera signals to e…
541207active
gali8/Tesseract-OCR-iOS
An iOS framework wrapping the Tesseract OCR engine (with Leptonica and image libraries) for use in Objective-C or Swift apps on iOS 9.0+. I…
234221maintenance
frotms/PaddleOCR2Pytorch
A PyTorch port of PaddleOCR that lets you run PaddleOCR-trained models (detection, recognition, and document structure parsing) without the…
731205active
dexsuite/dex-retargeting
A Python library of retargeting optimizers that translate human hand motion (from video or pose datasets) into robot dexterous hand joint c…
371205active
PoseLib/PoseLib
PoseLib is a C++ library of minimal solvers for calibrated camera pose estimation, covering absolute and relative pose from point and line …
681203active
dendenxu/fast-gaussian-rasterization
A drop-in replacement for diff-gaussian-rasterization that renders 3D Gaussian Splatting scenes using a geometry-shader-based GPU pipeline …
191202active
cleanlab/cleanvision
CleanVision is a Python library that automatically detects issues in image datasets, such as blurry, dark, over-exposed, or near-duplicate …
591199active
Calamari-OCR/calamari
Calamari is a Python-based OCR engine for line-based automatic text recognition, built on OCRopy and Kraken with a TensorFlow deep-learning…
741197active
toshas/torch-fidelity
A PyTorch library providing accurate and efficient implementations of generative model evaluation metrics such as FID, Inception Score, KID…
701197active
deepseek-ai/DeepSeek-VL
DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf…
254175maintenance
EvolvingLMMs-Lab/LLaVA-OneVision-2
A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis…
721195active
zai-org/CogAgent
CogAgent is an open-source vision-language model (VLM) based GUI agent that understands screen captures and natural language to automate in…
331194active
lessthanoptimal/BoofCV
BoofCV is an open-source, real-time computer vision library written entirely in Java, covering image processing, camera calibration, featur…
861192active
Tencent-Hunyuan/HunyuanWorld-Mirror
HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds…
541191active
trianglesplatting/triangle-splatting
Official implementation of 'Triangle Splatting for Real-Time Radiance Field Rendering' (3DV 2026), which uses 3D triangles as rendering pri…
431191active
shallowdream204/DreamClear
DreamClear is a diffusion-transformer based real-world image restoration model for high-fidelity super-resolution, published at NeurIPS 202…
271191active
zai-org/VisualGLM-6B
VisualGLM-6B is an open-source multimodal conversational language model supporting images, Chinese, and English, built on ChatGLM-6B with a…
304154maintenance
sair-lab/AirSLAM
AirSLAM is an efficient, illumination-robust point-line visual SLAM system supporting stereo visual odometry/VIO, offline map optimization,…
471190active
ali-vilab/UniAnimate
UniAnimate is the official code for a research paper on animating a reference human image into a video that follows a driving pose sequence…
311189active
pqpo/SmartCropper
An Android library for smart image cropping that automatically detects document borders using OpenCV (with an optional TensorFlow Lite HED …
664132maintenance
SHI-Labs/Neighborhood-Attention-Transformer
Official PyTorch implementation of the Neighborhood Attention Transformer (NAT/DiNAT), a family of efficient vision transformers with local…
321184stable
chensjtu/GaussianObject
GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,…
261183active
msracver/Deformable-ConvNets
Official MXNet implementation of Deformable Convolutional Networks (ICCV 2017) and R-FCN, including deformable convolution and ROI pooling …
324121maintenance
orpatashnik/StyleCLIP
Official implementation of StyleCLIP, a method for text-driven manipulation of StyleGAN-generated imagery using CLIP. It provides three app…
324121maintenance
VladimirYugay/Gaussian-SLAM
A research implementation of a dense RGBD SLAM system that uses 3D Gaussian Splatting as its scene representation to photorealistically rec…
271182active

← prev page 10 / 24 next →