Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
facebookresearch/multipathnet
A Torch-7 implementation of the MultiPath Network for object detection from the BMVC 2016 paper by Facebook AI Research, also supporting Fa…
101331abandoned
trailbehind/DeepOSM
DeepOSM is a Python application that trains neural networks with TensorFlow to classify roads and features in satellite imagery using OpenS…
321330abandoned
YonghaoHe/LFFD-A-Light-and-Fast-Face-Detector-for-Edge-Devices
LFFD is a light and fast single-class object detection framework designed for edge devices, with pretrained models for face, head, pedestri…
321323abandoned
datitran/object_detector_app
A Python application that performs real-time object recognition from a webcam or HLS video stream using TensorFlow's Object Detection API a…
321305abandoned
tianzhi0549/CTPN
Reference implementation of CTPN (Connectionist Text Proposal Network) for detecting text lines in natural images, from the ECCV 2016 paper…
321287abandoned
Prinsphield/Wechat_AutoJump
A Python bot that automatically plays the WeChat 'Jump Jump' mini-game using computer vision and a CNN coarse-to-fine model to locate the p…
321280abandoned
MicrosoftEdge/magic-mirror-demo
A smart mirror IoT demo project by Microsoft Edge that displays information on a two-way mirror and recognizes registered users via facial …
101225abandoned
CarlosGS/Cyclone-PCB-Factory
Cyclone PCB Factory is a parametric, 3D-printable CNC mill design (RepRap-style) intended for milling printed circuit boards. It provides C…
101187abandoned
Jai-wei/YOLOv8-PySide6-GUI
YoloSide is a desktop GUI application built with PySide6 for running YOLOv8 object detection models. Users can load trained .pt model files…
301171abandoned
linkedlist771/SoraWatermarkCleaner
A deep learning tool that detects and removes the Sora2 watermark from AI-generated videos using a YOLO-based detector plus a restoration m…
101149abandoned
faceair/youjumpijump
A Go-based cheat bot for the WeChat 'Jump Jump' (跳一跳) mini-game that screenshots the screen, computes jump distance via image analysis, and…
101146abandoned
eldar/pose-tensorflow
A TensorFlow implementation of the DeeperCut and ArtTrack algorithms for human body pose estimation, supporting both single-person and mult…
321141abandoned
szad670401/end-to-end-for-chinese-plate-recognition
An end-to-end Chinese license plate recognition model based on MXnet, using multi-label classification. It was trained on ~500k synthetic r…
321117abandoned
cgtinker/BlendArMocap
A Blender add-on that performs markerless motion capture using Google's Mediapipe, detecting pose, hand, and face features from webcam stre…
481078abandoned
mikebuss/MTBBarcodeScanner
A lightweight Objective-C barcode scanning library for iOS built on AVFoundation, supporting single and multiple barcode detection, torch c…
101075abandoned
tomthecarrot/arcore-for-all
A modified build of Google's ARCore developer preview library that removes the official device whitelist check, enabling ARCore to run on u…
101054abandoned
burningcl/wechat_jump_hack
A Java-based bot that automatically plays WeChat's 'Jump Jump' (跳一跳) mini-game by capturing screenshots via ADB, recognizing player and tar…
321053abandoned
facebookresearch/VMZ
VMZ is a model zoo from Facebook AI's Computer Vision team providing Caffe2 and PyTorch implementations of video classification models such…
101052abandoned
NVIDIA-AI-IOT/redtail
NVIDIA Redtail provides deep learning and computer vision components for autonomous visual navigation of drones and ground vehicles, center…
231047abandoned
digital-standard/ThreeDPoseTracker
A Unity-based Windows application that estimates 3D human pose from video or webcam input using an ONNX neural network model via Unity Barr…
231042abandoned
PRBonn/lidar-bonnetal
A deep learning framework for training and deploying semantic segmentation of LiDAR point clouds using range-image representations, develop…
101037abandoned
doug/depthjs
DepthJS is a browser extension and native plugin (primarily for Chrome) that lets any web page interact with the Microsoft Kinect via JavaS…
321001abandoned
tensorflow/models
The TensorFlow Model Garden is a repository of official and community implementations of state-of-the-art machine learning models built wit…
8577652active
deepfakes/faceswap
Faceswap is a free, open-source, multi-platform deepfakes tool that uses deep learning to recognize and swap faces in pictures and videos. …
8557500active
mudler/LocalAI
LocalAI is an open-source, self-hosted AI inference runtime that runs LLMs, vision, speech, image, and video models behind OpenAI/Anthropic…
9348696active
huggingface/pytorch-image-models
PyTorch Image Models (timm) is a Python library offering the largest collection of PyTorch image encoder/backbone architectures with 700+ p…
9337099active
78/xiaozhi-esp32
XiaoZhi is an open-source MCP-based AI voice chatbot firmware for ESP32-family microcontrollers, connecting large language models like Qwen…
8829186active
ente/ente
Ente is an open-source, end-to-end encrypted cloud platform with client apps for photos (a Google Photos alternative), document/credential …
9428513active
shap/shap
SHAP (SHapley Additive exPlanations) is a Python library that explains the output of any machine learning model using Shapley values from g…
9025704active
junyanz/pytorch-CycleGAN-and-pix2pix
Official PyTorch implementations of CycleGAN and pix2pix for paired and unpaired image-to-image translation. It includes training and testi…
4825232stable
deepseek-ai/DeepSeek-OCR
DeepSeek-OCR is an open vision-language model from DeepSeek AI that researches 'contexts optical compression' - encoding long text contexts…
4523855active
Skyvern-AI/skyvern
Skyvern is an open-source AI browser automation framework that uses LLMs and computer vision to interact with websites, offering a Playwrig…
8822852active
microsoft/unilm
Microsoft's collection of large-scale self-supervised pre-trained models spanning tasks, 100+ languages, and modalities (text, image, layou…
6722194active
huggingface/datasets
Hugging Face Datasets is a Python library providing one-line access to hundreds of thousands of public datasets on the Hugging Face Hub acr…
9821870stable
screenpipe/screenpipe
Screenpipe is a source-available desktop application that continuously records your screen and audio locally, extracting text via OCR/acces…
8621244active
huggingface/candle
Candle is a minimalist machine learning framework for Rust focused on performance and ease of use, with CPU and CUDA GPU support. It ships …
7320955active
QwenLM/Qwen3-VL
Qwen3-VL is a series of open-weight multimodal vision-language models from Alibaba's Qwen team, available in Dense and MoE architectures wi…
5219847active
huggingface/transformers.js
Transformers.js is a JavaScript library that lets you run Hugging Face Transformers pretrained models directly in the browser (or Node.js) …
9116270active
microsoft/Swin-Transformer
Official PyTorch implementation of the Swin Transformer, a hierarchical vision transformer using shifted windows that serves as a general-p…
3216051stable
alibaba/MNN
MNN is a lightweight, high-performance deep learning inference engine developed by Alibaba, supporting LLMs, vision, audio, and multimodal …
9315973active
duixcom/Duix-Avatar
Duix.Avatar is an open-source AI avatar toolkit for offline video generation and digital human cloning, capable of cloning a person's appea…
5714871active
dlib
Dlib is a modern C++ toolkit containing machine learning algorithms, deep learning tools, computer vision, linear algebra, and general-purp…
8614431stable
jacobgil/pytorch-grad-cam
A PyTorch library providing state-of-the-art pixel attribution (saliency) methods like GradCAM, ScoreCAM, and AblationCAM for explainable A…
7612958active
octalmage/robotjs
RobotJS is a Node.js desktop automation library for controlling the mouse and keyboard and reading the screen, with native prebuilt binarie…
9312770active
ludwig-ai/ludwig
Ludwig is a declarative, low-code deep learning framework for training, fine-tuning, and deploying AI models — from LLMs to tabular, image,…
9911745active
qubvel-org/segmentation_models.pytorch
A PyTorch library providing neural networks for image semantic segmentation with a simple high-level API. It includes 12 encoder-decoder ar…
7011706stable
milesial/Pytorch-UNet
A PyTorch implementation of the U-Net architecture for semantic segmentation of high-resolution images, originally built for Kaggle's Carva…
2311613active
facebookresearch/dinov3
Reference PyTorch implementation and pretrained models for DINOv3, Meta's self-supervised vision transformer backbone family. It includes t…
5911249active
voxel51/fiftyone
FiftyOne is an open-source Python library and GUI app for building high-quality computer vision datasets and models. It enables visualizing…
9911042active
OpenVINO
OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su…
9510740stable
autogluon/autogluon
AutoGluon is an AutoML library that automates machine learning on tabular data, time series, text, and images with just a few lines of Pyth…
9010617active
thumbor/thumbor
Thumbor is an open-source, on-demand image thumbnailing service written in Python. It crops, resizes, flips, and applies filters to images …
9010514active
openframeworks/openFrameworks
openFrameworks is an open-source C++ toolkit for creative coding that wraps common libraries like OpenGL, OpenCV, and audio/video libraries…
8710419stable
OpenGVLab/InternVL
InternVL is a family of open-source multimodal large language models (vision-language models) that combine vision transformers with LLMs to…
3710146active
open-mmlab/mmsegmentation
MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat…
239930stable
bytedance/Dolphin
Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-…
529049active
apple/ml-sharp
SHARP is a Python tool from Apple that synthesizes a photorealistic 3D Gaussian splat representation from a single photograph in under a se…
428843active
firerpa/lamda
FIRERPA (lamda) is an all-in-one Android device control platform whose server runs directly on the device (root or non-root) and exposes 16…
988243active
GetStream/Vision-Agents
An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (O…
848100active
facebookresearch/SlowFast
PySlowFast is a PyTorch-based open-source video understanding codebase from Facebook AI Research (FAIR). It provides implementations of sta…
657410active
farzaa/clicky
Clicky is an open-source macOS app that acts as an AI teacher living next to your cursor — it can see your screen, talk with you via voice,…
507394active
facebookresearch/sam-3d-objects
SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code…
557322active
turanszkij/WickedEngine
Wicked Engine is an open-source C++ 3D game engine with modern graphics features like ray tracing, global illumination, and physically base…
997203active
PaddlePaddle/models
PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda…
236932active
ml5js/ml5-library
ml5.js is a friendly, beginner-oriented JavaScript machine learning library for the browser, built on top of TensorFlow.js. It provides acc…
236587active
iperov/DeepFaceLive
DeepFaceLive is a real-time face-swap application for PC streaming and video calls, using trained face models (DFM) applied to webcam or vi…
1031011maintenance
PaddlePaddle/PaddleX
PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti…
926251active
om-ai-lab/VLM-R1
VLM-R1 is a framework for training R1-style large vision-language models using reinforcement learning (GRPO) on top of Qwen2.5-VL. It provi…
636015active
PaddlePaddle/PaddleClas
PaddleClas is a Python library and toolkit for image classification, recognition, and retrieval built on the PaddlePaddle deep learning fra…
665838active
aidlearning/AidLearning-FrameWork
AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling…
705797active
dnhkng/GLaDOS
A real-life implementation of GLaDOS, the sardonic AI from Valve's Portal series, built as a proactive voice assistant with vision, memory,…
645689active
microsoft/SynapseML
SynapseML (formerly MMLSpark) is an open-source machine learning library built on Apache Spark that provides simple, composable, distribute…
885240active
katanaml/sparrow
Sparrow is an open-source framework for structured data extraction from documents (PDFs, images) using ML, LLMs, and Vision LLMs, with sche…
955202active
ai-dawang/PlugNPlay-Modules
A curated collection of plug-and-play deep learning modules (convolutions, attention mechanisms, downsampling, and feature fusion blocks) i…
385105active
dmMaze/BallonsTranslator
A desktop GUI application that uses deep learning to automatically translate comics and manga, combining text detection, OCR, inpainting, a…
995065active
KaiyangZhou/deep-person-reid
Torchreid is a PyTorch library for deep-learning person re-identification, supporting both image and video reid with end-to-end training an…
504900stable
aloshdenny/reverse-SynthID
A research tool that reverse-engineers Google's SynthID watermark embedded in Gemini-generated images using spectral analysis and signal pr…
664815active
open-mmlab/mmocr
MMOCR is OpenMMLab's PyTorch-based toolbox for text detection, recognition, and key information extraction. It provides a model zoo of OCR …
234752active
Tencent/TNN
TNN is a high-performance, lightweight deep learning inference framework developed by Tencent Youtu Lab, supporting mobile, desktop, and se…
324648active
facebookresearch/vjepa2
Official PyTorch codebase and pretrained models for V-JEPA 2, a self-supervised video encoder trained on internet-scale video, plus V-JEPA …
524527active
layumi/Person_reID_baseline_pytorch
A small, friendly PyTorch baseline implementation for person and vehicle re-identification (ReID). It reproduces strong top-conference resu…
654446stable
xlite-dev/lite.ai.toolkit
A lightweight C++ toolkit providing unified APIs for 100+ pre-trained AI models across inference backends like ONNX Runtime, MNN, TensorRT,…
744427active
open-compass/VLMEvalKit
VLMEvalKit is an open-source Python toolkit for evaluating large vision-language models (LMMs/LVLMs) across 80+ benchmarks with support for…
624359active
httprunner/httprunner
HttpRunner (hrp) is an open-source, Go-based all-in-one testing framework for API testing (HTTP/HTTP2/WebSocket/RPC), load testing, and mul…
484295active
QwenLM/Qwen2.5-Omni
Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates stre…
314074active
thuml/Transfer-Learning-Library
TLlib is a PyTorch-based open-source library for transfer learning, covering domain adaptation, task adaptation (finetuning), and domain ge…
233931active
NVlabs/VILA
VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d…
573857active
open-mmlab/mmpretrain
MMPretrain is OpenMMLab's PyTorch-based toolbox and benchmark for image classification model pre-training, covering supervised, self-superv…
233850active
shitagaki-lab/see-through
A research framework from a SIGGRAPH 2026 paper that decomposes a single anime character illustration into up to 23 fully inpainted, semant…
583641active
NVlabs/Eagle
Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L…
643462active
Kedreamix/Linly-Talker
Linly-Talker is a Python-based digital avatar conversational system that combines LLMs (Linly, Qwen, Gemini), speech recognition (Whisper, …
483436active
opengeos/geoai
GeoAI is a Python package that integrates artificial intelligence with geospatial data analysis, built on PyTorch, Transformers, and segmen…
893327active
deepdoctection/deepdoctection
deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c…
983248active
Beckschen/TransUNet
Official PyTorch implementation of TransUNet, a U-Net-style architecture that uses a Vision Transformer encoder for medical image segmentat…
633234stable
facebookresearch/dinov2
PyTorch implementation and pretrained models for DINOv2, a self-supervised vision transformer method from Meta AI that learns robust visual…
6813266maintenance
junyanz/CycleGAN
A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I…
3212870maintenance
SharpAI/DeepCamera
DeepCamera is an open-source AI camera skills platform that runs local VLM scene analysis, object detection, face recognition, and person r…
863019active
osmr/imgclsmob
A research sandbox providing (re)implementations of numerous deep learning computer vision models for classification, segmentation, detecti…
233016active
Tamsiree/RxTool
RxTool is a large collection of utility classes and UI components for Android development, packaged as multiple Gradle modules (RxKit, RxUI…
2312288maintenance
sherlockchou86/VideoPipe
VideoPipe is a C++ framework for video analysis and structuring that works like a pipeline of independent, composable nodes. It integrates …
542931active

← prev page 13 / 16 next →