Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
BIT-DataLab/Edit-Banana
Edit Banana is an open-source Python framework that converts static images and PDFs of diagrams, flowcharts, and charts into fully editable…
605469active
xxlong0/Wonder3D
Wonder3D is a cross-domain diffusion model that reconstructs high-fidelity textured 3D meshes from a single image in 2-3 minutes. It genera…
325425active
isl-org/MiDaS
MiDaS is a Python library with pretrained models for robust monocular depth estimation from a single image, based on the TPAMI 2022 paper a…
105420stable
facebookresearch/sapiens
Sapiens is a family of foundation models from Meta Reality Labs for human-centric vision tasks including 2D pose estimation, body-part segm…
615418active
mayocream/koharu
Koharu is a local-first desktop application that automates manga translation using machine learning, combining text/bubble detection, OCR, …
825410active
deepseek-ai/DeepSeek-VL2
DeepSeek-VL2 is a series of Mixture-of-Experts vision-language models (Tiny, Small, and 4.5B activated parameters) with inference code and …
255374active
roboflow/sports
A Python library from Roboflow providing reusable computer vision tools for sports analytics, including ball tracking, player tracking and …
755320active
google-ar/arcore-android-sdk
Google's ARCore SDK for Android, providing Java and C APIs for building augmented reality experiences with motion tracking, environmental u…
795229active
katanaml/sparrow
Sparrow is an open-source framework for structured data extraction from documents (PDFs, images) using ML, LLMs, and Vision LLMs, with sche…
955202active
timesler/facenet-pytorch
A PyTorch library providing pretrained face detection (MTCNN) and facial recognition (Inception ResNet V1) models, ported from the TensorFl…
425162stable
NVIDIAGameWorks/kaolin
Kaolin is NVIDIA's PyTorch library of GPU-optimized modules for 3D deep learning research, covering meshes, point clouds, and 3D Gaussian s…
685161active
yisol/IDM-VTON
Official implementation of IDM-VTON, an ECCV 2024 paper that improves diffusion models for high-fidelity virtual try-on, swapping garments …
305156active
open-mmlab/mmaction2
MMAction2 is OpenMMLab's PyTorch-based toolbox and benchmark for video understanding, covering action recognition, temporal action localiza…
555142active
Breakthrough/PySceneDetect
PySceneDetect is a Python and OpenCV-based program and library for detecting scene cuts and transitions in videos, with multiple detection …
895123stable
ai-dawang/PlugNPlay-Modules
A curated collection of plug-and-play deep learning modules (convolutions, attention mechanisms, downsampling, and feature fusion blocks) i…
385105active
facebookresearch/AugLy
AugLy is a Python data augmentation library from Meta AI supporting audio, image, text, and video with over 100 augmentations. It focuses o…
675089stable
facebookresearch/co-tracker
CoTracker is a transformer-based model from Meta AI and Oxford VGG that jointly tracks any point (pixel) across a video, handling occlusion…
605080active
libigl/libigl
libigl is a simple C++ geometry processing library offering a wide range of algorithms for triangle and tetrahedral meshes, including defor…
675078stable
opentrack/opentrack
opentrack is a head tracking application that captures a user's head movements via webcams, IR trackers, or hardware devices and relays the…
755076active
dmMaze/BallonsTranslator
A desktop GUI application that uses deep learning to automatically translate comics and manga, combining text detection, OCR, inpainting, a…
995065active
Deci-AI/super-gradients
SuperGradients is an open-source PyTorch-based training library for building, training, and fine-tuning state-of-the-art computer vision mo…
545052active
ArthurBrussee/brush
Brush is a 3D reconstruction engine using Gaussian splatting, built in Rust on the Burn ML framework and WebGPU. It trains and renders spla…
654992active
KaiyangZhou/deep-person-reid
Torchreid is a PyTorch library for deep-learning person re-identification, supporting both image and video reid with end-to-end training an…
504900stable
tyxsspa/AnyText
AnyText is the official implementation of a diffusion-based model for multilingual visual text generation and editing in images, accepted a…
324874active
runhey/OnmyojiAutoScript
OnmyojiAutoScript (OAS) is a free, open-source automation script for the mobile game Onmyoji, built on the AzurLaneAutoScript framework. It…
654815active
UX-Decoder/Segment-Everything-Everywhere-All-At-Once
SEEM is the official PyTorch implementation of the NeurIPS 2023 paper 'Segment Everything Everywhere All at Once', a model for universal im…
204794stable
zju3dv/EasyMocap
EasyMocap is an open-source Python toolbox for markerless human motion capture and novel view synthesis from RGB videos. It fits parametric…
544783active
open-mmlab/mmocr
MMOCR is OpenMMLab's PyTorch-based toolbox for text detection, recognition, and key information extraction. It provides a model zoo of OCR …
234752active
OpenDriveLab/UniAD
UniAD is a unified end-to-end autonomous driving framework that hierarchically casts perception, prediction, and planning tasks under a pla…
444737active
cvg/LightGlue
LightGlue is a deep neural network library that matches sparse local features across image pairs with high accuracy and fast inference. It …
504728stable
esimov/pigo
Pigo is a pure Go library for fast face detection, pupil/eye localization, and facial landmark detection based on the Pixel Intensity Compa…
314728stable
MaaXYZ/MaaFramework
MaaFramework is an automation black-box testing framework based on image recognition, rewritten from the experience of the MAA (MaaAssistan…
914722active
manycore-research/SpatialLM
SpatialLM is a 3D large language model that processes point cloud data (from monocular video, RGBD images, or LiDAR) and generates structur…
624719active
CloudCompare/CloudCompare
CloudCompare is a 3D point cloud and triangular mesh processing application, originally built to compare point clouds from laser scanners a…
674693active
f3d-app/f3d
F3D is a fast, minimalist open-source 3D viewer desktop application supporting many formats (glTF, USD, STL, STEP, OBJ, FBX, Alembic) with …
884650active
Tencent/TNN
TNN is a high-performance, lightweight deep learning inference framework developed by Tencent Youtu Lab, supporting mobile, desktop, and se…
324648active
NVlabs/neuralangelo
Official PyTorch implementation of Neuralangelo, a CVPR 2023 method for high-fidelity neural surface reconstruction from multi-view images.…
294615active
sensity-ai/dot
dot (Deepfake Offensive Toolkit) is a Python tool that generates real-time, controllable deepfakes from a webcam feed and injects them into…
234586active
joanrod/star-vector
StarVector is a foundation model that generates scalable vector graphics (SVG) code from images and text by treating vectorization as a cod…
494560active
hku-mars/FAST-LIVO2
FAST-LIVO2 is a fast, tightly-coupled LiDAR-inertial-visual odometry and mapping system written in C++ on ROS. It provides real-time, accur…
564557active
ceres-solver/ceres-solver
Ceres Solver is an open-source C++ library for modeling and solving large-scale non-linear optimization problems, including bounded non-lin…
764547stable
facebookresearch/vjepa2
Official PyTorch codebase and pretrained models for V-JEPA 2, a self-supervised video encoder trained on internet-scale video, plus V-JEPA …
524527active
TencentARC/InstantMesh
InstantMesh is a feed-forward framework for generating 3D meshes from a single image using sparse-view large reconstruction models (LRM/Ins…
254509active
royshil/obs-backgroundremoval
An OBS Studio plugin that removes and replaces the background in portrait video using ONNX-based machine learning segmentation, acting as a…
984492active
layumi/Person_reID_baseline_pytorch
A small, friendly PyTorch baseline implementation for person and vehicle re-identification (ReID). It reproduces strong top-conference resu…
654446stable
rom1504/img2dataset
A Python tool that downloads large sets of image URLs and packages them into machine learning datasets, with resizing and caption support. …
564443active
xlite-dev/lite.ai.toolkit
A lightweight C++ toolkit providing unified APIs for 100+ pre-trained AI models across inference backends like ONNX Runtime, MNN, TensorRT,…
744427active
huawei-noah/Efficient-AI-Backbones
A collection of efficient neural network backbone architectures (GhostNet, TNT, ViG, WaveMLP, TinyNet, etc.) from Huawei Noah's Ark Lab, wi…
284418active
EvolvingLMMs-Lab/lmms-eval
lmms-eval is a unified Python framework for evaluating large multimodal models across text, image, video, and audio tasks, with 100+ benchm…
854377active
open-compass/VLMEvalKit
VLMEvalKit is an open-source Python toolkit for evaluating large vision-language models (LMMs/LVLMs) across 80+ benchmarks with support for…
624359active
dreamgaussian/dreamgaussian
DreamGaussian is the official PyTorch implementation of an ICLR 2024 Oral paper for efficient 3D content creation using generative Gaussian…
184352active
VectorSpaceLab/OmniGen
OmniGen is a unified diffusion-based image generation model that produces and edits images from multi-modal prompts without auxiliary modul…
474340active
unum-cloud/USearch
USearch is a fast, single-file similarity search and clustering engine for vectors and arbitrary objects, supporting spatial, binary, proba…
934278active
richzhang/PerceptualSimilarity
A PyTorch library implementing the LPIPS (Learned Perceptual Image Patch Similarity) metric, which measures perceptual distance between ima…
234269stable
SysCV/sam-hq
HQ-SAM (Segment Anything in High Quality) upgrades Meta's Segment Anything Model with a learnable High-Quality Output Token for accurate ze…
484255active
ali-vilab/AnyDoor
AnyDoor is the official implementation of a diffusion-based model that teleports target objects into new scenes at user-specified locations…
284238active
cvg/Hierarchical-Localization
hloc is a modular Python toolbox for state-of-the-art 6-DoF visual localization, combining image retrieval and feature matching (SuperPoint…
484194active
facebookresearch/vggt-omega
VGGT-Omega is a research library from Oxford VGG and Meta AI providing pretrained transformer models for 3D vision tasks such as camera pos…
584165active
torchgeo/torchgeo
TorchGeo is a PyTorch domain library, similar to torchvision, providing datasets, samplers, transforms, and pre-trained models specific to …
954159active
justadudewhohacks/face-api.js
A JavaScript face detection and face recognition library built on top of tensorflow.js, usable in the browser and Node.js. It provides mode…
2317945maintenance
solvespace/solvespace
SolveSpace is a free, open-source parametric 2D/3D CAD application with a constraint-based sketcher and solid modeling via extrudes, revolv…
804118active
WebODM/WebODM
WebODM is a user-friendly, commercial-grade application for drone image processing that generates georeferenced maps, point clouds, elevati…
984116active
VectorSpaceLab/OmniGen2
OmniGen2 is an open-source unified multimodal generation model supporting text-to-image generation, instruction-guided image editing, and i…
524112active
facebookresearch/jepa
Official PyTorch implementation of V-JEPA, a self-supervised method for learning visual representations from video using a joint-embedding …
294105active
cdcseacave/openMVS
OpenMVS is an open-source C++ library for Multi-View Stereo 3D reconstruction, taking camera poses and a sparse point-cloud as input and pr…
764100active
ZhengPeng7/BiRefNet
BiRefNet is a PyTorch implementation of the CAAI AIR 2024 paper 'Bilateral Reference for High-Resolution Dichotomous Image Segmentation'. I…
654098active
GuyTevet/motion-diffusion-model
Official PyTorch implementation of the Human Motion Diffusion Model (MDM) paper, generating 3D human motion sequences from text prompts usi…
524092active
princeton-vl/RAFT
Official PyTorch implementation of RAFT (Recurrent All Pairs Field Transforms for Optical Flow), an ECCV 2020 model for estimating dense op…
494091stable
StarsfieldAI/R1-V
R1-V is an open-source research codebase for training vision-language models with reinforcement learning (RLVR/GRPO), demonstrating strong …
214063active
Motion-Project/motion
Motion is an open-source C++ program that monitors video camera signals and detects changes (motion) in the images. It is commonly used for…
654038active
tensorflow/tensor2tensor
Tensor2Tensor (T2T) is a Python library of deep learning models and datasets built on TensorFlow, developed by the Google Brain team to mak…
1017464maintenance
cozmo/jsQR
jsQR is a pure JavaScript QR code reading library that takes raw image data (RGBA pixel arrays) and locates, extracts, and parses any QR co…
694026stable
introlab/rtabmap
RTAB-Map (Real-Time Appearance-Based Mapping) is a C++ library and standalone application implementing graph-based SLAM for RGB-D, stereo, …
883965stable
thuml/Transfer-Learning-Library
TLlib is a PyTorch-based open-source library for transfer learning, covering domain adaptation, task adaptation (finetuning), and domain ge…
233931active
hustvl/Vim
Vision Mamba (Vim) is a PyTorch implementation of a generic vision backbone built on bidirectional Mamba state space models, published at I…
293899active
hustvl/4DGaussians
An official PyTorch implementation of 4D Gaussian Splatting (4D-GS) for real-time rendering of dynamic scenes, published at CVPR 2024. It c…
273895active
tinyobjloader/tinyobjloader
A tiny, dependency-free Wavefront .obj/.mtl parser available as a single-header C++11 library and a pure C11 implementation with polygon te…
623866active
NVlabs/VILA
VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d…
573857active
mseitzer/pytorch-fid
A PyTorch port of the official TensorFlow implementation of the Fréchet Inception Distance (FID), a metric for measuring similarity between…
233851stable
open-mmlab/mmpretrain
MMPretrain is OpenMMLab's PyTorch-based toolbox and benchmark for image classification model pre-training, covering supervised, self-superv…
233850active
Avatarify
Avatarify is an open-source application that drives photorealistic avatars in real time for video-conferencing apps like Zoom and Skype, ba…
2316515maintenance
google-research/scenic
Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr…
763821active
CHNZYX/Auto_Simulated_Universe
A Python-based automation tool for the Honkai: Star Rail 'Simulated Universe' game mode, using screen recognition to play the roguelike mod…
873811active
lightly-ai/lightly
LightlySSL is a Python library built on PyTorch for self-supervised learning on images, offering modular implementations of methods like Si…
933797active
fudan-generative-vision/hallo2
Hallo2 is a Python research library from Fudan University that animates a single portrait image using audio input, producing long-duration …
263734active
abhiTronix/vidgear
VidGear is a high-performance, cross-platform Python framework for video processing built around multi-threaded and asynchronous pipelines.…
763721active
roboflow/trackers
A Python library of clean-room, Apache 2.0 implementations of multi-object tracking algorithms including SORT, ByteTrack, OC-SORT, BoT-SORT…
853717active
MaaEnd/MaaEnd
MaaEnd is a vision-AI-powered automation assistant for the game 'Arknights: Endfield', built on MaaFramework. It captures the screen, recog…
953712active
IDEA-Research/Grounded-SAM-2
Grounded SAM 2 is a foundation-model pipeline that combines open-set detectors (Grounding DINO, Grounding DINO 1.5/1.6, Florence-2, DINO-X)…
373708active
xinyu1205/recognize-anything
Recognize Anything is a collection of open-source image recognition foundation models, including RAM, RAM++, and Tag2Text, that perform ima…
333708active
Belval/TextRecognitionDataGenerator
A Python library and CLI tool (trdg) that generates synthetic text images for training OCR and text recognition models. It supports multipl…
233691stable
ant-research/MagicQuill
MagicQuill is an intelligent interactive image editing system from a CVPR 2025 paper, combining a brush-based UI with AI-powered suggestion…
463688active
microsoft/Bringing-Old-Photos-Back-to-Life
The official PyTorch implementation of 'Bringing Old Photos Back to Life' (CVPR 2020 Oral), a deep learning model that restores old photos …
2315704maintenance
DLR-RM/BlenderProc
BlenderProc is a procedural Python pipeline built on Blender for generating photorealistic synthetic training images with ground-truth anno…
623684active
mmp/pbrt-v4
pbrt-v4 is the C++ physically based ray tracing system accompanying the fourth edition of the book 'Physically Based Rendering: From Theory…
713683stable
stack-of-tasks/pinocchio
Pinocchio is a fast C++ library (with Python bindings) implementing state-of-the-art rigid body dynamics algorithms for poly-articulated sy…
933682active
facebookresearch/map-anything
MapAnything is an open-source research framework from Meta and CMU for universal feed-forward metric 3D reconstruction using an end-to-end …
773682active
ferdous-alam/GenCAD
GenCAD is a research codebase for image-conditioned CAD model generation using transformer-based contrastive representations (CCIP) and dif…
363669active
mikedh/trimesh
Trimesh is a pure Python library for loading, manipulating, and analyzing triangular meshes with an emphasis on watertight surfaces. It pro…
983660stable
borglab/gtsam
GTSAM is a C++ library implementing smoothing and mapping (SAM) for robotics and vision using factor graphs and Bayes networks as its core …
923656active

← prev page 3 / 24 next →