Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
open-mmlab/mmsegmentation
MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat…
239930stable
playcanvas/supersplat
SuperSplat is a free, open-source browser-based editor for inspecting, editing, optimizing, and publishing 3D Gaussian Splats. It is built …
909904active
StarTrail-org/PixelRAG
PixelRAG is a Python library and hosted service for visual retrieval-augmented generation: it renders web pages and documents into screensh…
769742active
mrousavy/react-native-vision-camera
A high-performance camera library for React Native offering photo/video capture, QR/barcode scanning, and JS worklet-based frame processors…
999580active
WongKinYiu/yolov9
Official PyTorch implementation of the YOLOv9 object detection paper, featuring Programmable Gradient Information for improved accuracy. It…
169551active
PeterL1n/RobustVideoMatting
Robust Video Matting (RVM) is a deep learning model and library for real-time human video matting, using a recurrent neural network with te…
239500stable
PaddlePaddle/PaddleSeg
PaddleSeg is an end-to-end image segmentation toolkit built on PaddlePaddle, offering a model zoo with dozens of pre-trained models for sem…
529382active
lipku/LiveTalking
LiveTalking is an open-source real-time interactive streaming digital human engine that drives a talking-head avatar from text or audio wit…
899238active
studio-dots-ai/dots.ocr
dots.ocr is a 1.7B-parameter vision-language model for multilingual document layout parsing, converting documents into structured output wi…
519090active
roboflow/rf-detr
RF-DETR is a real-time transformer-based model architecture from Roboflow for object detection, instance segmentation, and keypoint detecti…
879063active
bytedance/Dolphin
Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-…
529049active
RealSense SDK
RealSense SDK 2.0 (librealsense) is a cross-platform C++ library for Intel/RealSense depth cameras, providing depth and color streaming plu…
938974active
dusty-nv/jetson-inference
A C++/Python DNN inference library and tutorial guide ('Hello AI World') for deploying deep learning vision models on NVIDIA Jetson devices…
448969stable
apple/ml-sharp
SHARP is a Python tool from Apple that synthesizes a photorealistic 3D Gaussian splat representation from a single photograph in under a se…
428843active
PantsuDango/Dango-Translator
Dango-Translator (团子翻译器) is a Windows desktop application that performs real-time OCR-based translation of on-screen text ('raw' untranslat…
918751active
FoundationVision/VAR
Official PyTorch implementation of Visual Autoregressive Modeling (VAR), a NeurIPS 2024 Best Paper-winning method for scalable image genera…
488729active
DepthAnything/Depth-Anything-V2
Depth Anything V2 is a foundation model for monocular depth estimation, trained on 595K synthetic labeled images and 62M+ real unlabeled im…
568709stable
fudan-generative-vision/hallo
Hallo is a Python research library implementing hierarchical audio-driven visual synthesis for animating portrait images into talking-head …
148664active
jomjol/AI-on-the-edge-device
A firmware application for ESP32-CAM boards that uses TensorFlow Lite CNNs on-device to digitize analog utility meters (water, gas, electri…
728624active
CASIA-LMC-Lab/FastSAM
FastSAM is a CNN-based Segment Anything Model trained on only 2% of the SA-1B dataset, achieving comparable segmentation performance to SAM…
198401active
XPixelGroup/BasicSR
BasicSR is an open-source PyTorch toolbox for image and video restoration tasks such as super-resolution, denoising, deblurring, and JPEG a…
238367stable
bytedeco/javacv
JavaCV is a Java library that wraps OpenCV, FFmpeg, and other computer vision and multimedia libraries via JavaCPP Presets, with utility cl…
868335active
mikel-brostrom/boxmot
BoxMOT is a pluggable Python and C++ library providing state-of-the-art multi-object tracking (MOT) algorithms such as ByteTrack, BoT-SORT,…
958281active
exadel-inc/CompreFace
Exadel CompreFace is a free, open-source face recognition system that provides REST APIs for face recognition, verification, detection, lan…
238273stable
Ucas-HaoranWei/GOT-OCR2.0
Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, …
258216active
GetStream/Vision-Agents
An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O…
848100active
bingoogolapple/BGAQRCode-Android
An Android library for scanning and generating QR codes and barcodes, offering both ZXing and ZBar engines behind customizable scan views. …
648008stable
open-mmlab/mmpose
MMPose is an open-source pose estimation toolbox and benchmark built on PyTorch as part of the OpenMMLab ecosystem. It provides implementat…
397855active
wang-xinyu/tensorrtx
A C++ collection of popular deep learning networks (YOLO variants, ResNet, MobileNet, Swin Transformer, OCR models, and more) implemented f…
767827active
TencentARC/GFPGAN
GFPGAN is a Python library built on PyTorch that restores and enhances real-world degraded face photos using GAN-based priors. It provides …
2337657maintenance
TadasBaltrusaitis/OpenFace
OpenFace is a facial behavior analysis toolkit that performs facial landmark detection, head pose estimation, facial action unit recognitio…
237740active
boltgolt/howdy
Howdy provides Windows Hello-style facial authentication for Linux using IR emitters and a camera with facial recognition. It integrates wi…
387738active
geekyutao/Inpaint-Anything
Inpaint Anything combines Segment Anything (SAM) with inpainting models like LaMa and Stable Diffusion to remove, fill, or replace objects …
657703active
RapidAI/RapidOCR
RapidOCR is an open-source, multi-language OCR toolkit that performs text detection and recognition using models converted to run on ONNX R…
977599active
1adrianb/face-alignment
A Python library built on PyTorch that detects 2D and 3D facial landmarks in images using the FAN deep learning face alignment network. It …
707538active
hybridgroup/gocv
GoCV is a Go language binding for the OpenCV 4 computer vision library, supporting Linux, macOS, Windows, and Docker. It includes support f…
767491active
google/draco
Draco is an open-source C++ library from Google for compressing and decompressing 3D geometric meshes and point clouds. It improves the sto…
677458stable
apple/ml-fastvlm
Official implementation of FastVLM, a vision language model with an efficient hybrid vision encoder (FastViTHD) that reduces token count an…
287411active
facebookresearch/SlowFast
PySlowFast is a PyTorch-based open-source video understanding codebase from Facebook AI Research (FAIR). It provides implementations of sta…
657410active
zai-org/GLM-OCR
GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understa…
657366active
EutropicAI/Final2x
Final2x is a cross-platform desktop application for image super-resolution (upscaling) built with Electron, Vue3, and a PyTorch-based Pytho…
747323active
facebookresearch/sam-3d-objects
SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code…
557322active
naver/dust3r
DUSt3R is the official PyTorch implementation of a CVPR 2024 model that performs dense, unconstrained stereo and multi-view 3D reconstructi…
457288active
liuliu/ccv
ccv is a modern, minimalist computer vision library written in C/C++ with an application-driven set of state-of-the-art algorithms includin…
777243active
princeton-vl/infinigen
Infinigen is a procedural generator of infinite photorealistic 3D worlds and scenes, built on Blender by the Princeton Vision & Learning La…
947222active
BVLC/caffe
Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp…
2334556maintenance
PeterL1n/BackgroundMattingV2
Official PyTorch implementation of the CVPR 2021 paper 'Real-Time High-Resolution Background Matting'. It produces state-of-the-art alpha m…
237189stable
CMU-Perceptual-Computing-Lab/openpose
OpenPose is a real-time multi-person keypoint detection library that jointly estimates human body, face, hand, and foot keypoints (135 tota…
2334413maintenance
ok-oldking/ok-wuthering-waves
ok-ww is an image-recognition-based automation tool for the game Wuthering Waves, supporting background operation, automatic combat, echo f…
877167active
zxing/zxing
ZXing ('Zebra Crossing') is an open-source, multi-format 1D/2D barcode image processing library implemented in Java, with ports to other la…
7334077maintenance
yangchris11/samurai
SAMURAI is the official implementation of a zero-shot visual object tracker built on top of Segment Anything Model 2 (SAM 2), using a motio…
277112active
chenfei-wu/TaskMatrix
TaskMatrix (Visual ChatGPT) is a Python framework that connects ChatGPT with a suite of visual foundation models like Stable Diffusion, Gro…
3034003maintenance
OneDragon-Anything/ZenlessZoneZero-OneDragon
A Python-based automation assistant for the game Zenless Zone Zero that uses image recognition and OCR to fully automate daily tasks, dunge…
907029active
apple/corenet
CoreNet is Apple's deep neural network training toolkit for training standard and novel small and large-scale models, including foundation …
457007active
PaddlePaddle/models
PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda…
236932active
sczhou/ProPainter
ProPainter is a PyTorch-based video inpainting model from ICCV 2023 that combines dual-domain propagation with a mask-guided sparse video T…
226916stable
VAST-AI-Research/TripoSR
TripoSR is an open-source model for fast feedforward 3D object reconstruction from a single image, developed by Tripo AI and Stability AI. …
646888active
FoundationVision/ByteTrack
ByteTrack is a PyTorch-based multi-object tracking (MOT) library implementing the ECCV 2022 paper 'Multi-Object Tracking by Associating Eve…
326654stable
Yuliang-Liu/MonkeyOCR
MonkeyOCR is a lightweight large multimodal model (LMM) for document parsing that uses a Structure-Recognition-Relation triplet paradigm to…
606635active
scikit-image/scikit-image
scikit-image is a Python library providing a collection of peer-reviewed image processing algorithms built on NumPy and SciPy. It offers ro…
826577stable
iperov/DeepFaceLive
DeepFaceLive is a real-time face-swap application for PC streaming and video calls, using trained face models (DFM) applied to webcam or vi…
1031011maintenance
openMVG/openMVG
OpenMVG is a C++ library for multiple view geometry and Structure from Motion (SfM), providing end-to-end 3D reconstruction from images. It…
486542active
AILab-CVC/YOLO-World
YOLO-World is a real-time open-vocabulary object detection model and Python toolkit from Tencent AI Lab and HUST, published at CVPR 2024. I…
296529active
open-mmlab/mmdetection3d
MMDetection3D is OpenMMLab's next-generation platform for general 3D object detection, built on PyTorch. It provides a modular toolbox with…
236518active
open-mmlab/mmcv
MMCV is the foundational computer vision library for the OpenMMLab ecosystem, providing image/video I/O, data transformations, and CUDA ope…
526470stable
sz3/libcimbar
libcimbar is an optimized C++ implementation of the cimbar (color icon matrix) high-density 2D barcode format for air-gapped data transfer …
956426active
OpenDroneMap/ODM
OpenDroneMap (ODM) is an open source command line toolkit that processes aerial drone, balloon, or kite imagery into classified point cloud…
906417active
madmaze/pytesseract
Python-tesseract is a Python wrapper for Google's Tesseract-OCR engine that recognizes and extracts text embedded in images. It supports al…
646383stable
KevinMusgrave/pytorch-metric-learning
A PyTorch library providing modular losses, miners, samplers, and testers for deep metric learning. It supports building complete train/tes…
486339active
mindee/doctr
docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe…
906315active
szad670401/HyperLPR
HyperLPR3 is a high-performance open-source framework for recognizing Chinese license plates, built with deep learning and available as a P…
276255active
RangiLyu/nanodet
NanoDet-Plus is a super fast, lightweight anchor-free object detection model implemented in PyTorch, with model sizes as small as 980KB (IN…
236252stable
PaddlePaddle/PaddleX
PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti…
926251active
ByteDance-Seed/Depth-Anything-3
Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,…
596213active
ByteDance-Seed/Bagel
BAGEL is an open-source unified multimodal foundation model with 7B active parameters (14B total) trained on interleaved multimodal data. I…
556159active
open-edge-platform/anomalib
Anomalib is a Python deep learning library for anomaly detection, offering state-of-the-art unsupervised algorithms for detecting and local…
986088active
shimat/opencvsharp
OpenCvSharp is a cross-platform .NET wrapper for the OpenCV computer vision library, published as NuGet packages with bundled native binari…
986072active
bytedance/LatentSync
LatentSync is an end-to-end lip-sync framework from ByteDance based on audio-conditioned latent diffusion models, using Stable Diffusion to…
336026active
CGAL/cgal
CGAL (Computational Geometry Algorithms Library) is a mature C++ library providing efficient and reliable algorithms for computational geom…
956024stable
om-ai-lab/VLM-R1
VLM-R1 is a framework for training R1-style large vision-language models using reinforcement learning (GRPO) on top of Qwen2.5-VL. It provi…
636015active
Doubiiu/ToonCrafter
ToonCrafter is a generative model that interpolates two cartoon images into a short animation by leveraging pre-trained image-to-video diff…
296003stable
OFA-Sys/Chinese-CLIP
Chinese-CLIP is a Chinese version of the CLIP model trained on ~200 million Chinese image-text pairs, built on open_clip. It provides APIs,…
665998active
journeyapps/zxing-android-embedded
An Android barcode scanning library built on the ZXing decoder, usable via Intents or embedded in an Activity for custom UI. It supports po…
235931stable
ChaoningZhang/MobileSAM
MobileSAM is the official implementation of a lightweight version of Meta's Segment Anything Model (SAM), replacing the heavyweight image e…
655858stable
PaddlePaddle/PaddleClas
PaddleClas is a Python library and toolkit for image classification, recognition, and retrieval built on the PaddlePaddle deep learning fra…
665838active
cnr-isti-vclab/meshlab
MeshLab is an open source, portable, and extensible system for processing and editing unstructured large 3D triangular meshes, built on the…
675802active
pjreddie/darknet
Darknet is an open-source neural network framework written in C and CUDA, best known as the original home of the YOLO real-time object dete…
3226492maintenance
DeepLabCut/DeepLabCut
DeepLabCut is an open-source Python toolbox for markerless 2D and 3D pose estimation of user-defined body parts using deep neural networks …
895745stable
NVIDIA/DALI
NVIDIA DALI is a GPU-accelerated data loading and preprocessing library with optimized building blocks and an execution engine for deep lea…
925734active
ladaapp/lada
Lada is an open-source tool with both GUI and CLI that restores pixelated/mosaic regions in videos, primarily targeting JAV (Japanese adult…
655702active
open-mmlab/OpenPCDet
OpenPCDet is a PyTorch-based open-source toolbox for LiDAR-based 3D object detection. It provides official implementations of models like P…
535692active
apple/ml-depth-pro
Depth Pro is Apple's reference implementation of a foundation model for zero-shot metric monocular depth estimation, producing sharp high-r…
305683active
idealo/imagededup
imagededup is a Python library for finding exact and near-duplicate images in a collection using perceptual hashing algorithms (PHash, DHas…
485666stable
Fanghua-Yu/SUPIR
SUPIR is a Python-based photo-realistic image restoration system built on SDXL diffusion priors and LLaVA captioning, presented at CVPR 202…
365649active
Vision-CAIR/MiniGPT-4
Official code for MiniGPT-4 and MiniGPT-v2, vision-language models that align a frozen visual encoder with a frozen LLM (Vicuna) via a sing…
3025627maintenance
nerfstudio-project/gsplat
gsplat is an open-source Python library with CUDA-accelerated, differentiable rasterization of Gaussians, based on 3D Gaussian Splatting fo…
705589active
matterport/Mask_RCNN
A Python implementation of the Mask R-CNN model for object detection and instance segmentation, built on Keras and TensorFlow with a ResNet…
2325567maintenance
ngxson/smolvlm-realtime-webcam
A browser-based demo that streams webcam frames to a llama.cpp server running SmolVLM 500M for real-time object detection and scene descrip…
295571active
lyuwenyu/RT-DETR
Official implementation of RT-DETR and RT-DETRv2, real-time object detection transformers that outperform YOLO models, in PyTorch and Paddl…
745476active
obss/sahi
SAHI (Slicing Aided Hyper Inference) is a Python vision library for detecting small objects in large images via sliced/tiled inference, wor…
995473active

← prev page 2 / 24 next →