Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
shitagaki-lab/see-through
A research framework from a SIGGRAPH 2026 paper that decomposes a single anime character illustration into up to 23 fully inpainted, semant…
583641active
cmusatyalab/openface
OpenFace is a free and open source Python and Torch implementation of face recognition based on Google's FaceNet deep neural network. It ge…
6515438maintenance
facebookresearch/detr
DETR is Facebook Research's PyTorch implementation of Detection Transformer, an end-to-end object detection model that replaces hand-crafte…
1015354maintenance
albumentations-team/albumentations
Albumentations is a fast, flexible Python image augmentation library for computer vision, supporting images, masks, bounding boxes, keypoin…
1015315maintenance
MrNeRF/LichtFeld-Studio
LichtFeld Studio is a native open-source desktop application for 3D Gaussian Splatting that combines training, real-time inspection, splat …
923594active
ZhaoJ9014/face.evoLVe
A high-performance face recognition library built on PaddlePaddle and PyTorch, providing comprehensive tools for face-related analytics and…
383589active
AiuniAI/Unique3D
Unique3D is the official implementation of a NeurIPS 2024 paper that generates high-quality textured 3D meshes from a single image in about…
383579active
SkalskiP/make-sense
makesense.ai is a free, browser-based tool for labeling photos to prepare datasets for computer vision projects. It runs entirely client-si…
233562active
GVCLab/PersonaLive
PersonaLive is a diffusion-based framework for real-time, streamable portrait image animation, generating infinite-length expressive talkin…
533552active
ToTheBeginning/PuLID
PuLID is the official PyTorch implementation of a NeurIPS 2024 method for inserting a specific person's identity into text-to-image generat…
403550active
AliaksandrSiarohin/first-order-model
Official PyTorch/Jupyter implementation of the First Order Motion Model for image animation (NeurIPS 2019). It animates a static source ima…
3215015maintenance
sparkjsdev/spark
Spark is an advanced 3D Gaussian Splatting renderer library for THREE.js, built in TypeScript by World Labs. It integrates with the THREE.j…
813529active
google-research/big_vision
Google Research's official Jax/Flax codebase for training large-scale vision models such as Vision Transformer, SigLIP, MLP-Mixer, and LiT …
423528active
cszn/KAIR
A PyTorch image restoration toolbox providing training and testing code for many restoration models including DnCNN, FFDNet, SRMD, USRNet, …
233523active
MashiroSaber03/Saber-Translator
Saber-Translator is an AI-powered manga translation application that detects speech bubbles, OCRs Japanese text, translates it, inpaints th…
783516active
NVlabs/FoundationPose
FoundationPose is NVIDIA's unified foundation model for 6D object pose estimation and tracking of novel objects, supporting both model-base…
623516active
MooreThreads/Moore-AnimateAnyone
An open-source reproduction of AnimateAnyone that animates a character from a single reference image using pose sequences from a driving vi…
263514active
visionml/pytracking
PyTracking is a PyTorch-based framework for visual object tracking and video object segmentation, providing official implementations of tra…
233514active
PKU-YuanGroup/Video-LLaVA
Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into…
273500active
Anttwo/SuGaR
SuGaR is the official PyTorch implementation of a CVPR 2024 method that extracts accurate, editable meshes from 3D Gaussian Splatting recon…
273495active
aleju/imgaug
imgaug is a Python library for augmenting images in machine learning experiments, converting a small set of input images into a much larger…
2314741maintenance
facebookresearch/ijepa
Official PyTorch implementation of I-JEPA, a self-supervised learning method that predicts latent representations of image regions from oth…
103489active
AliceVision
Meshroom is an open-source, node-based visual programming application for building and executing data processing pipelines, best known for …
693486active
vietanhdev/anylabeling
AnyLabeling is a desktop image annotation tool that combines LabelImg/Labelme-style manual labeling with AI-assisted auto-labeling. It runs…
853463active
NVlabs/Eagle
Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L…
643462active
RainerKuemmerle/g2o
g2o is an open-source C++ framework for optimizing graph-based nonlinear error functions, commonly used for nonlinear least squares problem…
673461stable
facebookresearch/sam-3d-body
SAM 3D Body is a promptable model for single-image full-body 3D human mesh recovery (HMR), estimating body, feet, and hand pose using the M…
483461active
nihui/waifu2x-ncnn-vulkan
A portable command-line tool implementing the waifu2x anime-style image upscaler and denoiser using the ncnn inference framework with the V…
633456active
roflcoopter/viseron
Viseron is a self-hosted, local-only network video recorder (NVR) with built-in AI computer vision capabilities. It supports object detecti…
983431active
davidsandberg/facenet
A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC…
3214343maintenance
luigifreda/pyslam
pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras with a wide range of local and global fe…
773401active
IQA-PyTorch
A pure Python/PyTorch toolbox for image quality assessment (IQA) providing GPU-accelerated reimplementations of many full-reference and no-…
823380active
deepseek-ai/DeepSeek-OCR-2
DeepSeek-OCR 2 is an open-source vision-language model and inference toolkit implementing 'Visual Causal Flow' for optical character recogn…
443379active
WongKinYiu/yolov7
Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques…
2314139maintenance
jenly1314/ZXingLite
ZXingLite is a streamlined, fast Android library built on ZXing for scanning and generating QR codes and barcodes, with fully customizable …
823367active
OpenTalker/SadTalker
SadTalker is a CVPR 2023 deep learning tool that generates realistic talking head videos from a single portrait image and an audio clip by …
2214040maintenance
VainF/Torch-Pruning
Torch-Pruning is a PyTorch framework for structural neural network pruning based on the DepGraph algorithm from CVPR 2023. It automatically…
503348active
OpenGVLab/Ask-Anything
VideoChat/Ask-Anything is a family of multimodal chat models and demos that combine video understanding with large language models, letting…
723346active
nihui/opencv-mobile
opencv-mobile provides minimal, prebuilt OpenCV binary packages for Android, iOS, ARM Linux, Windows, Linux, macOS, HarmonyOS, WebAssembly,…
853345active
RKNN-Toolkit2
RKNN-Toolkit2 is Rockchip's SDK for converting trained neural network models into RKNN format and deploying them on Rockchip NPU chips like…
363313active
Peterande/D-FINE
D-FINE is the official PyTorch implementation of an ICLR 2025 Spotlight paper that redefines the regression task in DETR-style detectors as…
673305active
XiaoMi/xiaomi-miloco
Xiaomi Miloco is an open-source whole-home intelligence solution that uses Mi Home camera video/audio as a full-modal perception gateway an…
823292active
hbb1/2d-gaussian-splatting
Official implementation of 2D Gaussian Splatting (2DGS), a SIGGRAPH 2024 method that represents scenes as 2D oriented Gaussian disks for ge…
693279stable
jixiaozhong/Sonic
Sonic is the official PyTorch implementation of the CVPR 2025 paper 'Sonic: Shifting Focus to Global Audio Perception in Portrait Animation…
493273active
vladmandic/human
Human is a JavaScript/TypeScript library built on TensorFlow.js that combines multiple ML models for 3D face detection and recognition, bod…
483264active
deepdoctection/deepdoctection
deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c…
983248active
Beckschen/TransUNet
Official PyTorch implementation of TransUNet, a U-Net-style architecture that uses a Vision Transformer encoder for medical image segmentat…
633234stable
mit-han-lab/bevfusion
BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re…
103230stable
Jittor/jittor
Jittor is a high-performance deep learning framework from Tsinghua University based on just-in-time (JIT) compilation and meta-operators, w…
673229active
breezedeus/Pix2Text
Pix2Text is an open-source Python tool that recognizes layouts, tables, math formulas (LaTeX), and text in images and converts them into Ma…
993227active
MzeroMiko/VMamba
VMamba is a PyTorch implementation of a visual state space model (SSM) vision backbone based on Mamba, featuring 2D Selective Scan (SS2D) f…
213219active
kerlomz/captcha_trainer
A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren…
553213active
OpenGVLab/InternGPT
InternGPT (iGPT) is an open-source demo platform for showcasing AI models through a pointing-language-driven visual interactive system, sup…
293205active
VTK
VTK (Visualization Toolkit) is an open-source C++ library for 3D graphics, image processing, volume rendering, and scientific visualization…
773201stable
PJLab-ADG/SensorsCalibration
OpenCalib is a C++ multi-sensor calibration toolbox for autonomous driving that calibrates IMU, LiDAR, camera, and radar sensors, both intr…
323200active
xianfei/SysMocap
SysMocap is a cross-platform, video-driven real-time motion capture system that animates 3D virtual characters from webcam footage. It rend…
783199active
prs-eth/Marigold
Marigold is a family of diffusion-based models and a fine-tuning protocol that adapts pretrained latent diffusion models like Stable Diffus…
523198active
Pointcept
Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan…
763196active
facebookresearch/dinov2
PyTorch implementation and pretrained models for DINOv2, a self-supervised vision transformer method from Meta AI that learns robust visual…
6813266maintenance
GuidoBartoli/sherloq
Sherloq is an open-source digital image forensic toolset providing an integrated GUI environment for analyzing images for tampering and aut…
743193active
ARM-software/ComputeLibrary
Arm's Compute Library is a C++ collection of over 100 low-level machine learning and computer vision functions optimized for Arm Cortex-A/N…
963183active
Rudrabha/Wav2Lip
Wav2Lip is the official research code for the ACM Multimedia 2020 paper 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In …
4513182maintenance
rmurai0610/MASt3R-SLAM
MASt3R-SLAM is a real-time monocular dense SLAM system built on the MASt3R two-view 3D reconstruction prior, producing globally consistent …
433167active
TurixAI/TuriX-CUA
TuriX is an open-source computer-use agent (CUA) that lets AI models take real actions on a desktop GUI - clicking, typing, and navigating …
703156active
cleardusk/3DDFA_V2
3DDFA_V2 is the official PyTorch implementation of the ECCV 2020 paper 'Towards Fast, Accurate and Stable 3D Dense Face Alignment'. It regr…
233149stable
megvii-research/NAFNet
NAFNet is the official PyTorch implementation of a state-of-the-art image restoration network that removes nonlinear activation functions. …
323148stable
automeris-io/WebPlotDigitizer
WebPlotDigitizer is a computer vision assisted web application that extracts numerical data from images of charts and plots. It has been wi…
653146active
z-x-yang/Segment-and-Track-Anything
An open-source pipeline (SAM-Track) that segments and tracks arbitrary objects in videos using the Segment Anything Model for key-frame seg…
613134active
otiai10/gosseract
gosseract is a Go package that provides OCR (Optical Character Recognition) by binding to the Tesseract C++ library via cgo. It lets Go app…
513130active
junyanz/CycleGAN
A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I…
3212870maintenance
jina-ai/clip-as-service
CLIP-as-service is a low-latency, high-scalability server for embedding images and text into fixed-length vectors using OpenAI's CLIP model…
2312836maintenance
naver/mast3r
MASt3R is the official PyTorch implementation of 'Grounding Image Matching in 3D with MASt3R' (ECCV 2024), a model that performs dense 3D r…
373088active
zxingify/zxingify-objc
ZXingObjC is a full Objective-C port of the ZXing barcode image processing library, supporting encoding and decoding of many 1D and 2D barc…
233075active
antimatter15/splat
A WebGL-based real-time viewer for 3D Gaussian Splatting scenes, rendering photorealistic navigable 3D environments from photo-derived spla…
513065active
AnyListen/tools-ocr
Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model…
233065active
rpng/open_vins
OpenVINS is an open-source C++ platform for visual-inertial navigation research, centered on a filter-based (MSCKF/EKF) estimator that fuse…
473057active
yuyuyzl/EasyVtuber
EasyVtuber is a Python-based VTubing application built on the Talking Head Anime model that turns a single anime character illustration int…
623051active
thiagoalessio/tesseract-ocr-for-php
A PHP wrapper library around the Tesseract OCR command-line binary, providing a fluent API for extracting text from images. It supports mul…
613040stable
SharpAI/DeepCamera
DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r…
863019active
osmr/imgclsmob
A research sandbox providing (re)implementations of numerous deep learning computer vision models for classification, segmentation, detecti…
233016active
ZQPei/deep_sort_pytorch
A PyTorch implementation of the Deep SORT multi-object tracking algorithm, pairing YOLOv3/YOLOv5 (or Mask R-CNN) detectors with a CNN re-id…
323012active
Doubiiu/DynamiCrafter
DynamiCrafter is an open-source research model that animates open-domain still images into short videos using pre-trained video diffusion p…
273007active
williamyang1991/Rerender_A_Video
The official PyTorch implementation of 'Rerender A Video', a SIGGRAPH Asia 2023 zero-shot text-guided video-to-video translation framework.…
292999stable
MeiGen-AI/MultiTalk
MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a …
562992active
iscyy/ultralyticsPro
A PyTorch-based collection of improved YOLO-family object detection models (YOLOv5 through YOLOv13, RT-DETR) with pluggable modules for bac…
482954active
zju3dv/LoFTR
LoFTR is a detector-free local image feature matching method using Transformers, released with PyTorch inference and training code plus pre…
322950stable
sunsmarterjie/yolov12
YOLOv12 is a PyTorch implementation of attention-centric real-time object detectors, published at NeurIPS 2025. It provides detection model…
592947active
microsoft/table-transformer
Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un…
232939active
KichangKim/DeepDanbooru
DeepDanbooru is a Python/TensorFlow system that estimates Danbooru-style tags for anime-style girl images using multi-label classification.…
632937active
sherlockchou86/VideoPipe
VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates …
542931active
jeeliz/jeelizFaceFilter
A lightweight JavaScript/WebGL library for real-time face detection and tracking from a camera feed via WebRTC, designed for building augme…
462928active
InternLM/InternLM-XComposer
InternLM-XComposer is a family of large vision-language models from the InternLM team, including XComposer2.5 for long-context multimodal u…
382925active
xelatihy/yocto-gl
Yocto/GL is a collection of small C++17 libraries for building physically-based graphics algorithms, written in a data-oriented style and r…
232925active
ogkalu2/comic-translate
An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language…
902911active
mitsuba-renderer/mitsuba3
Mitsuba 3 is a research-oriented, retargetable rendering system for forward and inverse light transport simulation, written in C++17 on top…
942899active
learnables/learn2learn
learn2learn is a PyTorch library for meta-learning research, providing utilities for few-shot task creation, high-level wrappers for algori…
482893active
nimiq/qr-scanner
A lightweight JavaScript/TypeScript QR code scanner library based on Cosmo Wolfe's port of Google's ZXing library. It supports webcam video…
322886stable
NVlabs/FoundationStereo
FoundationStereo is NVIDIA's official PyTorch implementation of a foundation model for zero-shot stereo depth estimation, published as a CV…
472874active
UX-Decoder/Semantic-SAM
Official PyTorch implementation of Semantic-SAM, a universal image segmentation model that segments and recognizes anything at any desired …
332854active
openmv/openmv
OpenMV is an open-source machine vision platform consisting of camera hardware firmware programmable in Python 3 (MicroPython). The firmwar…
912850active

← prev page 4 / 24 next →