Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: computer-vision

2316 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
hacksider/Deep-Live-Cam
Deep-Live-Cam is a Python application that performs real-time face swapping on webcam feeds and one-click video deepfakes using only a sing…
8996140active
ruvnet/RuView
RuView is a WiFi sensing platform that uses Channel State Information (CSI) from commodity WiFi hardware like ESP32 to detect presence, tra…
8191756active
OpenCV
OpenCV is the de facto open-source computer vision library, providing thousands of optimized algorithms for image and video processing, fea…
8990613stable
PaddlePaddle/PaddleOCR
PaddleOCR is a multilingual OCR and document parsing toolkit built on PaddlePaddle that converts images and PDFs into structured data like …
9388312stable
tensorflow/models
The TensorFlow Model Garden is a repository of official and community implementations of state-of-the-art machine learning models built wit…
8577652active
Tesseract OCR
Tesseract is an open-source OCR engine consisting of the libtesseract library and a command-line program, using an LSTM-based neural networ…
8676200stable
ultralytics/ultralytics
Ultralytics YOLO is a Python package and CLI providing a family of real-time computer vision models (YOLO26, YOLO11, YOLOv8) for object det…
9560991active
facebookresearch/segment-anything
Segment Anything Model (SAM) from Meta AI is a promptable image segmentation foundation model that produces high-quality object masks from …
3054759stable
roboflow/supervision
Supervision is a Python library of reusable computer vision tools that bridges the gap between detection/segmentation/classification models…
9549745active
hiroi-sora/Umi-OCR
Umi-OCR is a free, open-source, fully offline OCR application for Windows and Linux with a Qt/QML GUI. It supports screenshot OCR, batch im…
4746882stable
naptha/tesseract.js
Tesseract.js is a pure JavaScript port of the Tesseract OCR engine that extracts text from images in over 100 languages. It runs in the bro…
7038671active
huggingface/pytorch-image-models
PyTorch Image Models (timm) is a Python library offering the largest collection of PyTorch image encoder/backbone architectures with 700+ p…
9337099active
google-ai-edge/mediapipe
MediaPipe is Google's cross-platform framework for deploying on-device machine learning solutions for live and streaming media. It provides…
9436731stable
Real-ESRGAN
Real-ESRGAN is a deep learning project for practical image and video restoration via super-resolution, with pretrained models for photos an…
2336593stable
XingangPan/DragGAN
Official PyTorch implementation of DragGAN (SIGGRAPH 2023), an interactive point-based image manipulation method built on StyleGAN3. Users …
2935755stable
Frigate
Frigate is an open-source, self-hosted network video recorder (NVR) that performs real-time AI object detection on IP camera feeds locally …
9235407active
facebookresearch/detectron2
Detectron2 is Facebook AI Research's PyTorch-based library for state-of-the-art object detection, instance/panoptic segmentation, and other…
6734688stable
openai/CLIP
OpenAI's CLIP is a PyTorch library providing pretrained contrastive language-image models that encode images and text into a shared embeddi…
6634236stable
FreeCAD/FreeCAD
FreeCAD is a free and open-source, cross-platform 3D parametric CAD modeler for designing real-life objects of any size, built on the OpenC…
8933083stable
open-mmlab/mmdetection
MMDetection is OpenMMLab's PyTorch-based toolbox and benchmark for object detection and instance/panoptic segmentation. It provides a large…
2332892stable
JaidedAI/EasyOCR
EasyOCR is a ready-to-use Python OCR library built on PyTorch that extracts text from images, supporting 80+ languages and popular writing …
4829942stable
deepinsight/insightface
InsightFace is an open-source 2D and 3D face analysis project providing state-of-the-art face detection, recognition, alignment, and face s…
6529580active
Label Studio
Label Studio is an open-source data labeling and annotation platform supporting images, audio, text, video, and time series with a web UI a…
8628150active
OpenBMB/MiniCPM-V
MiniCPM-V and MiniCPM-o are a series of small multimodal large language models for efficient image, video, and audio understanding, deploya…
6126240active
lucidrains/vit-pytorch
A PyTorch library implementing the Vision Transformer (ViT) and dozens of ViT variants (NaViT, MaxViT, MobileViT, Dino, masked autoencoders…
8725488active
microsoft/OmniParser
OmniParser is a screen parsing tool from Microsoft that converts UI screenshots into structured, understandable elements to ground vision-l…
6025310active
junyanz/pytorch-CycleGAN-and-pix2pix
Official PyTorch implementations of CycleGAN and pix2pix for paired and unpaired image-to-image translation. It includes training and testi…
4825232stable
haotian-liu/LLaVA
LLaVA (Large Language and Vision Assistant) is an open-source multimodal large language model framework implementing visual instruction tun…
2025000active
baidu/Unlimited-OCR
Baidu's Unlimited-OCR is an open vision-language OCR model for one-shot long-horizon document parsing, extending DeepSeek-OCR. It provides …
5624569active
danielgatis/rembg
Rembg is a Python tool for removing image backgrounds using U2Net-based deep learning models. It can be used as a CLI, Python library, HTTP…
9824449active
graphdeco-inria/gaussian-splatting
The official reference implementation of 3D Gaussian Splatting, a method for real-time radiance field rendering that reconstructs scenes fr…
5023425active
serengil/deepface
DeepFace is a lightweight Python library for face recognition and facial attribute analysis, wrapping state-of-the-art models like VGG-Face…
8923340stable
MAA (MaaAssistantArknights)
MAA (MAA Assistant Arknights) is a C++ desktop assistant for the mobile game Arknights that automates daily tasks using image recognition. …
9422796active
huggingface/datasets
Hugging Face Datasets is a Python library providing one-line access to hundreds of thousands of public datasets on the Hugging Face Hub acr…
9821870stable
Zeyi-Lin/HivisionIDPhotos
HivisionIDPhotos is a lightweight AI tool that generates standard ID/passport photos from user images using offline matting models that run…
7121420active
datalab-to/surya
Surya is a 650M parameter OCR toolkit from Datalab providing state-of-the-art text recognition, layout analysis, reading order detection, a…
8621318active
bloc97/Anime4K
Anime4K is a set of open-source, high-quality real-time anime upscaling and denoising algorithms implemented as GLSL shaders, primarily for…
2321295stable
QwenLM/Qwen3-VL
Qwen3-VL is a series of open-weight multimodal vision-language models from Alibaba's Qwen team, available in Dense and MoE architectures wi…
5219847active
facebookresearch/sam2
Official code for Meta's Segment Anything Model 2 (SAM 2), a foundation model for promptable visual segmentation in images and videos. It i…
6119770active
KlingAIResearch/LivePortrait
LivePortrait is a Python-based portrait animation tool from Kuaishou Technology that synthesizes lifelike videos from a single source image…
6218969active
sczhou/CodeFormer
CodeFormer is a PyTorch-based blind face restoration model using a codebook lookup transformer, published at NeurIPS 2022. It restores and …
4618117stable
pytorch/vision
torchvision is the official PyTorch companion library providing datasets, model architectures, and image/video transformations for computer…
9317885stable
deepseek-ai/Janus
Janus-Series is DeepSeek's family of unified multimodal models (Janus, Janus-Pro, JanusFlow) that combine multimodal understanding and imag…
2417757active
IDEA-Research/Grounded-Segment-Anything
Grounded-Segment-Anything (Grounded SAM) combines Grounding DINO with Segment Anything to detect and segment arbitrary objects from text pr…
3017710active
NVlabs/instant-ngp
NVIDIA's implementation of instant neural graphics primitives, training NeRFs, signed distance functions, neural images, and neural volumes…
5217535stable
Robbyant/lingbot-map
LingBot-Map is a feed-forward 3D foundation model that reconstructs scenes from streaming image data using a Geometric Context Transformer.…
5816705active
cvat-ai/cvat
CVAT (Computer Vision Annotation Tool) is an open-source, self-hosted platform for annotating images, videos, and 3D point clouds to build …
9516600active
lukas-blecher/LaTeX-OCR
pix2tex (LaTeX-OCR) is a PyTorch-based vision transformer model that converts images of math formulas into LaTeX code. It ships as a pip-in…
2416547stable
huggingface/transformers.js
Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) …
9116270active
wkentaro/labelme
Labelme is a graphical image annotation tool written in Python with a Qt interface, supporting polygon, rectangle, oriented rectangle, circ…
9916130active
microsoft/Swin-Transformer
Official PyTorch implementation of the Swin Transformer, a hierarchical vision transformer using shifted windows that serves as a general-p…
3216051stable
babalae/better-genshin-impact
BetterGI is a free, open-source Windows desktop application that automates gameplay in Genshin Impact using computer vision, OCR, and YOLO-…
9515067active
duixcom/Duix-Avatar
Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea…
5714871active
tensorflow/tfjs-models
A collection of pre-trained machine learning models ported to TensorFlow.js, published as npm packages for use in JavaScript projects. Mode…
6314793active
HumanAIGC/AnimateAnyone
Animate Anyone is the official research implementation of a diffusion-based image-to-video synthesis method that animates a static characte…
4614791active
ddddocr
DdddOcr is a Python library for offline, local recognition of various CAPTCHA types, including alphanumeric, Chinese character, and slider …
6414665active
dlib
Dlib is a modern C++ toolkit containing machine learning algorithms, deep learning tools, computer vision, linear algebra, and general-purp…
8614431stable
PaddlePaddle/PaddleDetection
PaddleDetection is an object detection toolkit built on the PaddlePaddle deep learning framework. It provides implementations of detection,…
7314389active
facebookresearch/vggt
VGGT (Visual Geometry Grounded Transformer) is a feed-forward transformer model from Meta AI and Oxford VGG that infers 3D geometry—camera …
5814292active
mlfoundations/open_clip
OpenCLIP is an open-source PyTorch implementation of CLIP and related multimodal contrastive models, with many pretrained image/text checkp…
8614095active
img2threejs/img2threejs
A tool that reconstructs objects from reference images as code-only, procedural Three.js models rather than meshes or photogrammetry. It pr…
8014018active
Open3D
Open3D is an open-source C++ and Python library for 3D data processing, offering data structures, algorithms, and pipelines for point cloud…
6713913active
jacobgil/pytorch-grad-cam
A PyTorch library providing state-of-the-art pixel attribution (saliency) methods like GradCAM, ScoreCAM, and AblationCAM for explainable A…
7612958active
jwagner/smartcrop.js
smartcrop.js is a JavaScript library that implements a content-aware algorithm to find good crops for images. It runs in the browser, in No…
2312955stable
ShiqiYu/libfacedetection
An open-source C++ library for CNN-based face detection in images, with the model embedded as static C source so it has no external depende…
6312784stable
google-research/vision_transformer
Google Research's official JAX/Flax implementation of Vision Transformer (ViT) and MLP-Mixer architectures, with released pretrained checkp…
7512683stable
colmap/colmap
COLMAP is a general-purpose Structure-from-Motion (SfM) and Multi-View Stereo (MVS) pipeline for reconstructing 3D models from ordered or u…
9812564active
YaoFANGUK/video-subtitle-remover
An AI-based desktop application that removes hard-coded subtitles and text-like watermarks from videos and images using deep learning inpai…
7112553active
DayBreak-u/chineseocr_lite
An ultra-lightweight Chinese OCR toolkit combining DBNet text detection, CRNN text recognition, and an angle classifier, with total model s…
7012339active
simular-ai/Agent-S
Agent S is an open-source agentic framework that uses multimodal LLMs to operate computers like a human, controlling GUIs via clicking, typ…
7012193active
xmu-xiaoma666/External-Attention-pytorch
A PyTorch library (fightingcv-attention) providing clean, minimal implementations of numerous attention mechanisms, MLP variants, re-parame…
6512183active
datalab-to/chandra
Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO…
7112171active
instantX-research/InstantID
InstantID is a tuning-free, zero-shot identity-preserving image generation method built on diffusion models, generating customized images i…
2611987active
nerfstudio-project/nerfstudio
Nerfstudio is a Python library and CLI toolkit providing a simple, modular API for creating, training, and testing Neural Radiance Fields (…
3811934active
qubvel-org/segmentation_models.pytorch
A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar…
7011706stable
milesial/Pytorch-UNet
A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva…
2311613active
facebookresearch/sam3
Official code for Meta's Segment Anything Model 3 (SAM 3), a unified foundation model for promptable segmentation in images and videos. It …
6311487active
UI-TARS
UI-TARS is ByteDance's open-source multimodal AI agent stack, comprising Agent TARS (a CLI/Web UI multimodal agent that operates terminals,…
4911389active
rerun-io/rerun
Rerun is an open-source SDK and viewer for logging, storing, querying, and visualizing multi-rate multimodal data such as images, point clo…
9911362active
THU-MIG/yolov10
YOLOv10 is a real-time end-to-end object detection model family that removes NMS post-processing via consistent dual assignments and optimi…
2011336active
kornia/kornia
Kornia is a differentiable computer vision library built on PyTorch, offering GPU-accelerated image processing, augmentations, and geometri…
8611327active
salesforce/LAVIS
LAVIS is a Python library from Salesforce AI Research providing a unified toolkit for language-vision (multimodal) intelligence, including …
6111262active
facebookresearch/dinov3
Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t…
5911249active
ultralytics/yolov5
Ultralytics YOLOv5 is a PyTorch-based computer vision model family for real-time object detection, instance segmentation, and image classif…
6757929maintenance
PointCloudLibrary/pcl
The Point Cloud Library (PCL) is a large-scale, modular open-source C++ library for 2D/3D image and point cloud processing. It provides sta…
6811101stable
voxel51/fiftyone
FiftyOne is an open-source Python library and GUI app for building high-quality computer vision datasets and models. It enables visualizing…
9911042active
ageitgey/face_recognition
A Python library and command-line tool providing a simple API for face detection, facial landmark extraction, and face recognition, built o…
6356684maintenance
microsoft/TRELLIS.2
TRELLIS.2 is a 4B-parameter state-of-the-art 3D generative model from Microsoft for high-fidelity image-to-3D asset generation. It uses a f…
5710869active
OpenVINO
OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su…
9510740stable
autogluon/autogluon
AutoGluon is an AutoML library that automates machine learning on tabular data, time series, text, and images with just a few lines of Pyth…
9010617active
Megvii-BaseDetection/YOLOX
YOLOX is a high-performance anchor-free YOLO object detection model family implemented in PyTorch, with pretrained weights and export suppo…
3410587stable
IDEA-Research/GroundingDINO
Official PyTorch implementation of Grounding DINO, an open-set object detector that fuses a Transformer-based DINO detector with grounded v…
2110515stable
esimov/caire
Caire is a content-aware image resize library written in Go, based on the seam carving algorithm. It intelligently shrinks or enlarges imag…
3110465active
openframeworks/openFrameworks
openFrameworks is an open-source C++ toolkit for creative coding that wraps common libraries like OpenGL, OpenCV, and audio/video libraries…
8710419stable
zyddnys/manga-image-translator
A Python tool that automatically detects, OCRs, translates, inpaints, and re-typesets text in images, primarily for manga and comics. It ru…
6510345active
CVHub520/X-AnyLabeling
X-AnyLabeling is a cross-platform desktop application for AI-assisted annotation of text, image, video, and multimodal data. It bundles bui…
9610212active
OpenGVLab/InternVL
InternVL is a family of open-source multimodal large language models (vision-language models) that combine vision transformers with LLMs to…
3710146active
freemocap/freemocap
FreeMoCap is a free, open-source, markerless motion capture system that uses ordinary cameras (webcams, GoPros, smartphones) to record and …
9810085active
m87-labs/moondream
Moondream is an open-weight family of small, efficient vision language models (2B to 9B MoE) that perform image captioning, visual question…
6110014active
facebookresearch/pytorch3d
PyTorch3D is Facebook AI Research's library of efficient, reusable components for deep learning with 3D data, built on PyTorch. It provides…
749954active

page 1 / 24 next →