Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: computer-vision

1555 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
qualcomm/ai-hub-models
Qualcomm AI Hub Models is a curated collection of 300+ state-of-the-art machine learning models (vision, audio, speech, generative AI) pre-…
891195active
mlmed/torchxrayvision
TorchXRayVision is an open-source PyTorch library providing pre-trained deep learning models and a unified interface for publicly available…
911183active
chensjtu/GaussianObject
GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,…
261183active
mlfoundations/open_flamingo
OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima…
234118maintenance
DAMO-NLP-SG/VideoLLaMA3
VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de…
371179active
JDAI-CV/fast-reid
FastReID is a PyTorch-based research platform implementing state-of-the-art re-identification algorithms for persons, vehicles, and faces. …
233981maintenance
smthemex/ComfyUI_Sonic
A ComfyUI custom node implementing the Sonic method for audio-driven portrait animation, generating talking-head videos from a single portr…
561140active
TowhidKashem/snapchat-clone
A Snapchat clone web application built with React, Redux Toolkit, and TypeScript, featuring camera-based face filters with Three.js augment…
641136active
espressif/esp-dl
ESP-DL is Espressif's lightweight neural network inference framework for ESP-series chips, with a custom .espdl model format, quantization …
721125active
caiyuanhao1998/MST
A Python toolbox for spectral compressive imaging reconstruction that implements over 15 algorithms including MST, CST, DAUHST, BiSCI, HDNe…
531118active
FutureUniant/Tailor
Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea…
371115active
gabber-dev/gabber
Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr…
441111active
szymanowiczs/splatter-image
Official PyTorch implementation of 'Splatter Image: Ultra-Fast Single-View 3D Reconstruction' (CVPR 2024), which uses an image-to-image net…
261106active
WangLibo1995/GeoSeg
GeoSeg is an open-source PyTorch-based semantic segmentation toolbox focused on Vision Transformers for remote sensing imagery, featuring t…
321096active
SCLBD/DeepfakeBench
DeepfakeBench is a comprehensive benchmark framework for deepfake detection, providing a unified platform for data management, implementati…
361093active
louiszengCN/CarlaAir
CarlaAir is an open-source simulation infrastructure that combines CARLA's high-fidelity urban driving environments with physics-accurate m…
741076active
autonomousvision/navsim
NAVSIM is a data-driven pseudo-simulation framework and benchmark for autonomous vehicle planning, evaluating driving agents non-reactively…
501076active
jeffbass/imagezmq
imageZMQ is a set of Python classes that transport OpenCV images between computers using PyZMQ messaging. It enables distributed computer v…
621069stable
InternRobotics/InternNav
InternNav is an open-source PyTorch-based toolbox for building embodied navigation foundation models, supporting vision-language navigation…
581061active
brenpoly/be-more-agent
An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp…
641060active
open-gigaai/giga-models
GigaModels is an open-source Python framework providing pipelines for training, inference, deployment, and compression of multi-modal, gene…
621057active
InternRobotics/PointLLM
PointLLM is a multimodal large language model that understands colored 3D point clouds of objects, built on a point cloud encoder fused wit…
651051active
liuyuan-pal/SyncDreamer
SyncDreamer is a synchronized multiview diffusion model that generates multiview-consistent images from a single-view image, released with …
501045active
Xiaoqi-Zhao-DLUT/MSNet-M2SNet
Official PyTorch implementations of MSNet and M2SNet, multi-scale subtraction networks for medical image segmentation such as polyp, lung i…
741042active
videoflow/videoflow
Videoflow is a Python framework for building distributed video and stream processing pipelines as directed acyclic graphs of producers, pro…
661034active
podgorskiy/ALAE
Official PyTorch implementation of Adversarial Latent Autoencoders (ALAE/StyleALAE), a CVPR 2020 paper combining autoencoders with GAN trai…
323511maintenance
tensorlayer/SRGAN
Reference implementation of SRGAN, a generative adversarial network for photo-realistic single image super-resolution, built on TensorLayer…
233466maintenance
abizovnuralem/go2_ros2_sdk
An unofficial ROS2 SDK for the Unitree Go2 quadruped robot (AIR/PRO/EDU), connecting over WebRTC (Wi-Fi) or CycloneDDS (Ethernet). It provi…
571020active
bowang-lab/U-Mamba
U-Mamba is a hybrid CNN-state-space-model (Mamba) network for biomedical image segmentation, built on top of the nnU-Net framework. It comb…
261019active
shaoanlu/faceswap-GAN
A Jupyter Notebook-based implementation of face swapping using a denoising autoencoder architecture enhanced with adversarial losses, VGGFa…
323416maintenance
meetps/pytorch-semseg
A PyTorch library implementing popular semantic segmentation architectures such as FCN, U-Net, SegNet, PSPNet, ICNet, FRRN, and LinkNet, wi…
233402maintenance
huawei-noah/noah-research
A collection of research code subprojects released by Huawei Noah's Ark Lab, each in its own directory. It is not an official Huawei produc…
761004active
siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System
A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro…
391001active
catalyst-team/catalyst
Catalyst is a high-level PyTorch framework for deep learning research and development, focused on reproducibility, rapid experimentation, a…
643382maintenance
aserbao/AndroidCamera
An Android library and demo app implementing a TikTok-style custom camera with video and audio editing features such as segment recording, …
233297maintenance
zhanghang1989/ResNeSt
ResNeSt is a PyTorch implementation of the Split-Attention Network, a ResNet variant that applies channel-wise attention across network bra…
233261maintenance
jacksonliam/mjpg-streamer
A lightweight command-line streaming tool that copies JPEG frames from input plugins (webcams, Raspberry Pi camera, files, OpenCV) to outpu…
323247maintenance
Tramac/awesome-semantic-segmentation-pytorch
A PyTorch library providing concise, modifiable reference implementations of many semantic segmentation models such as FCN, PSPNet, DeepLab…
323069maintenance
cvlab-columbia/zero123
Zero-1-to-3 is a research codebase and pretrained diffusion model from Columbia CVLab that changes the camera viewpoint of an object from a…
303058maintenance
Kurento/kurento-media-server
Kurento Media Server is a C++/GStreamer-based media server that handles media transmission, processing, recording, and streaming over WebRT…
103053maintenance
jfzhang95/pytorch-deeplab-xception
A PyTorch implementation of the DeepLab v3+ semantic segmentation model with support for multiple backbones (Xception, ResNet, MobileNet, D…
323000maintenance
tensorflow/graphics
TensorFlow Graphics is a library of differentiable graphics layers for TensorFlow, including differentiable renderers, spatial transformers…
642781maintenance
VainF/DeepLabV3Plus-Pytorch
A PyTorch library providing pretrained DeepLabv3 and DeepLabv3+ semantic segmentation models for Pascal VOC and Cityscapes datasets. It inc…
322699maintenance
yerfor/GeneFace
GeneFace is the official PyTorch implementation of an ICLR 2023 paper on generalized, high-fidelity audio-driven 3D talking face synthesis …
212657maintenance
HypoX64/DeepMosaics
DeepMosaics is a Python application that automatically removes or adds mosaics in images and videos using semantic segmentation and image-t…
232633maintenance
ShawnBIT/UNet-family
A curated collection of UNet-family semantic segmentation models with PyTorch implementations and links to original papers and third-party …
322592maintenance
CASIA-IVA-Lab/DANet
DANet is the official PyTorch implementation of 'Dual Attention Network for Scene Segmentation' (CVPR 2019), which uses position and channe…
322463maintenance
zai-org/CogVLM2
CogVLM2 is an open-source multi-modal vision-language model family built on Meta-Llama-3-8B-Instruct, offering image and video understandin…
282433maintenance
iPERDance/iPERCore
Impersonator++ (iPERCore) is a PyTorch implementation of Liquid Warping GAN with Attention, a unified framework for human image synthesis. …
322393maintenance
nbei/Deep-Flow-Guided-Video-Inpainting
A PyTorch implementation of the CVPR 2019 paper 'Deep Flow-Guided Video Inpainting', which fills missing regions in videos by completing op…
322375maintenance
michuanhaohao/reid-strong-baseline
A PyTorch implementation of the 'Bag of Tricks and A Strong Baseline for Deep Person Re-identification' paper (CVPRW 2019), providing end-t…
322355maintenance
hzwer/ICCV2019-LearningToPaint
A PyTorch research implementation of the ICCV 2019 paper 'Learning to Paint With Model-based Deep Reinforcement Learning'. It trains agents…
412305maintenance
donnyyou/torchcv
TorchCV is a PyTorch-based framework providing reimplementations of deep learning models for major computer vision tasks. It covers image c…
322251maintenance
bigmb/Unet-Segmentation-Pytorch-Nest-of-Unets
A PyTorch implementation of several U-Net variants for image segmentation, including UNet, R2U-Net, Attention U-Net, Attention R2U-Net, and…
322249maintenance
idealo/image-quality-assessment
A Python implementation of Google's NIMA (Neural Image Assessment) models that predict the aesthetic and technical quality of images using …
102243maintenance
atriumlts/subpixel
A TensorFlow reimplementation of the efficient sub-pixel convolutional neural network (ESPCN) for single-image super-resolution, based on S…
322123maintenance
qubvel/efficientnet
A Keras and TensorFlow Keras reimplementation of the EfficientNet convolutional neural network family (B0-B7), including ImageNet-pretraine…
232100maintenance
ozan-oktay/Attention-Gated-Networks
A PyTorch implementation of attention gates for convolutional neural networks, applied to U-Net and VGG-16 architectures. It targets medica…
322064maintenance
apple/ml-fastvit
Official PyTorch implementation of FastViT, a fast hybrid vision transformer architecture using structural reparameterization, published at…
282027maintenance
JDAI-CV/FaceX-Zoo
FaceX-Zoo is a PyTorch toolbox for face recognition that provides training modules with various state-of-the-art supervisory heads and back…
321999maintenance
albertpumarola/GANimation
Official PyTorch implementation of GANimation, an ECCV'18 research paper that animates facial expressions in a single image using a GAN con…
321985maintenance
apple/ml-cvnets
CVNets is Apple's open-source PyTorch library for training computer vision networks, covering classification, detection, segmentation, vide…
321983maintenance
WuJie1010/Facial-Expression-Recognition.Pytorch
A PyTorch implementation of CNN-based facial expression recognition achieving state-of-the-art accuracy on FER2013 (73.112%) and CK+ (94.64…
321976maintenance
SummitKwan/transparent_latent_gan
TL-GAN is a Python/TensorFlow project that makes a GAN's latent space transparent by discovering feature axes, enabling controlled image sy…
321973maintenance
ronghuaiyang/arcface-pytorch
A PyTorch implementation of ArcFace, a deep metric learning approach for face recognition that adds angular margin penalties to face embedd…
321901maintenance
TreB1eN/InsightFace_Pytorch
A PyTorch reimplementation of InsightFace/ArcFace for face recognition, including backbone models (IR-SE50, MobileFacenet) and pretrained w…
321894maintenance
open-mmlab/mmaction
MMAction is an open-source PyTorch toolbox for video action understanding, covering action recognition, temporal action detection, and spat…
321876maintenance
NVIDIA/semantic-segmentation
NVIDIA's PyTorch monorepo implementing the paper 'Hierarchical Multi-Scale Attention for Semantic Segmentation', with pretrained weights an…
321828maintenance
yassouali/pytorch-segmentation
A PyTorch library implementing multiple semantic segmentation models (DeepLab V3+, PSPNet, U-Net, SegNet, FCN, ENet, and others) with datas…
261818maintenance
MCG-NJU/VideoMAE
Official PyTorch implementation of VideoMAE, a masked autoencoder method for data-efficient self-supervised video pre-training with video t…
321784maintenance
cszn/DnCNN
DnCNN is the official implementation of the TIP 2017 paper 'Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising',…
321727maintenance
bubbliiiing/unet-pytorch
A PyTorch implementation of the U-Net semantic segmentation model with training, prediction, and mIoU evaluation scripts. It supports multi…
231725maintenance
svip-lab/impersonator
A PyTorch implementation of Liquid Gating GAN (ICCV 2019) that performs human motion imitation, appearance transfer, and novel view synthes…
321717maintenance
dunbar12138/pix2pix3D
pix2pix3D is the official PyTorch implementation of a CVPR 2023 paper on 3D-aware conditional image synthesis. It generates 3D objects (neu…
311716maintenance
Temporal Segment Networks (TSN)
Official code and pretrained models for Temporal Segment Networks (TSN), a deep learning framework for video action recognition published a…
321577maintenance
HumanAIGC/EMO
EMO (Emote Portrait Alive) is a research codebase from Alibaba's Institute for Intelligent Computing that generates expressive talking port…
257594experimental
sniklaus/3d-ken-burns
A PyTorch reference implementation of the 3D Ken Burns Effect from a Single Image paper, which animates a still photo with a virtual camera…
701569maintenance
AlexHex7/Non-local_pytorch
A PyTorch implementation of the Non-local Neural Block from the paper 'Non-local Neural Networks', providing multiple variants (concatenati…
321564maintenance
open-mmlab/Multimodal-GPT
Multimodal-GPT is an open-source project for training a multimodal chatbot that combines vision and language instructions, built on OpenFla…
301512maintenance
szagoruyko/attention-transfer
PyTorch reference implementation of the ICLR 2017 paper 'Paying More Attention to Attention', which improves convolutional neural networks …
321463maintenance
datitran/face2face-demo
A pix2pix demo that learns from facial landmarks and translates them into a trained target face, with a real-time webcam application. It in…
321461maintenance
facebookresearch/MaskFormer
MaskFormer is a PyTorch/Detectron2-based implementation of the NeurIPS 2021 paper 'Per-Pixel Classification is Not All You Need for Semanti…
101460maintenance
hysts/pytorch_image_classification
A PyTorch library implementing many image classification architectures (ResNet, DenseNet, WRN, PyramidNet, SENet, etc.) and augmentation te…
101449maintenance
megvii-research/ML-GCN
A PyTorch implementation of ML-GCN, the CVPR 2019 paper 'Multi-Label Image Recognition with Graph Convolutional Networks'. It provides trai…
321446maintenance
BachiLi/redner
redner is a differentiable Monte Carlo ray tracer that computes exact gradients of rendered images with respect to arbitrary scene paramete…
321444maintenance
xuebinqin/BASNet
BASNet is the official PyTorch implementation of the CVPR 2019 paper 'BASNet: Boundary-Aware Salient Object Detection', a deep learning mod…
321436maintenance
vlfeat/matconvnet
MatConvNet is a MATLAB toolbox implementing convolutional neural networks (CNNs) for computer vision applications. It supports training and…
321430maintenance
rmokady/CLIP_prefix_caption
Official implementation of ClipCap, a CLIP-based image captioning model that maps CLIP image encodings to a GPT-2 prefix to generate captio…
321423maintenance
gradslam/gradslam
gradslam is a fully differentiable dense SLAM library built on PyTorch, providing differentiable building blocks such as nonlinear least sq…
231422maintenance
yu-changqian/TorchSeg
A fast, modular PyTorch reference implementation for training and evaluating semantic segmentation models such as FCN, DFN, BiSeNet, PSPNet…
231410maintenance
open-mmlab/mmfashion
MMFashion is an open-source PyTorch-based toolbox for visual fashion analysis from the OpenMMLab project. It provides modular implementatio…
321369maintenance
orobix/retina-unet
A Python implementation of a U-Net convolutional neural network for segmenting blood vessels in retina fundus images. It performs binary pi…
321354maintenance
aitorzip/PyTorch-CycleGAN
A clean, readable PyTorch implementation of CycleGAN for unpaired image-to-image translation using cycle-consistent adversarial networks. I…
321319maintenance
hukkelas/DeepPrivacy
DeepPrivacy is a PyTorch-based GAN that automatically anonymizes faces in images and videos by generating realistic synthetic replacements.…
321316maintenance
d-li14/involution
Official PyTorch implementation of the involution neural operator from the CVPR 2021 paper 'Involution: Inverting the Inherence of Convolut…
321310maintenance
bubbliiiing/deeplabv3-plus-pytorch
A PyTorch implementation of the DeepLabv3+ semantic segmentation model with MobileNetV2 and Xception backbones. It includes scripts for tra…
231292maintenance
YuanxunLu/LiveSpeechPortraits
A PyTorch implementation of the SIGGRAPH Asia 2021 paper 'Live Speech Portraits', which generates photorealistic personalized talking-head …
321283maintenance
elanmart/cbp-translate
A demo application that live-translates foreign-language speech in videos into subtitles, mimicking the Cyberpunk 2077 translation effect. …
321274maintenance
marcbelmont/cnn-watermark-removal
A TensorFlow implementation of a fully convolutional neural network that removes transparent watermark overlays from images. It trains on s…
321269maintenance
DrSleep/tensorflow-deeplab-resnet
A TensorFlow re-implementation of the DeepLab-ResNet model for semantic image segmentation, trained and evaluated on the PASCAL VOC dataset…
101258maintenance

← prev page 15 / 16 next →