function: computer-vision
1555 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| qualcomm/ai-hub-models Qualcomm AI Hub Models is a curated collection of 300+ state-of-the-art machine learning models (vision, audio, speech, generative AI) pre-… | 89 | 1195 | active |
| mlmed/torchxrayvision TorchXRayVision is an open-source PyTorch library providing pre-trained deep learning models and a unified interface for publicly available… | 91 | 1183 | active |
| chensjtu/GaussianObject GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,… | 26 | 1183 | active |
| mlfoundations/open_flamingo OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima… | 23 | 4118 | maintenance |
| DAMO-NLP-SG/VideoLLaMA3 VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de… | 37 | 1179 | active |
| JDAI-CV/fast-reid FastReID is a PyTorch-based research platform implementing state-of-the-art re-identification algorithms for persons, vehicles, and faces. … | 23 | 3981 | maintenance |
| smthemex/ComfyUI_Sonic A ComfyUI custom node implementing the Sonic method for audio-driven portrait animation, generating talking-head videos from a single portr… | 56 | 1140 | active |
| TowhidKashem/snapchat-clone A Snapchat clone web application built with React, Redux Toolkit, and TypeScript, featuring camera-based face filters with Three.js augment… | 64 | 1136 | active |
| espressif/esp-dl ESP-DL is Espressif's lightweight neural network inference framework for ESP-series chips, with a custom .espdl model format, quantization … | 72 | 1125 | active |
| caiyuanhao1998/MST A Python toolbox for spectral compressive imaging reconstruction that implements over 15 algorithms including MST, CST, DAUHST, BiSCI, HDNe… | 53 | 1118 | active |
| FutureUniant/Tailor Tailor is an AI-powered desktop video editing application offering intelligent video cutting, generation, and optimization. It provides fea… | 37 | 1115 | active |
| gabber-dev/gabber Gabber is an open-source engine for building real-time multimodal AI applications that can see, hear, and speak, using graph-based orchestr… | 44 | 1111 | active |
| szymanowiczs/splatter-image Official PyTorch implementation of 'Splatter Image: Ultra-Fast Single-View 3D Reconstruction' (CVPR 2024), which uses an image-to-image net… | 26 | 1106 | active |
| WangLibo1995/GeoSeg GeoSeg is an open-source PyTorch-based semantic segmentation toolbox focused on Vision Transformers for remote sensing imagery, featuring t… | 32 | 1096 | active |
| SCLBD/DeepfakeBench DeepfakeBench is a comprehensive benchmark framework for deepfake detection, providing a unified platform for data management, implementati… | 36 | 1093 | active |
| louiszengCN/CarlaAir CarlaAir is an open-source simulation infrastructure that combines CARLA's high-fidelity urban driving environments with physics-accurate m… | 74 | 1076 | active |
| autonomousvision/navsim NAVSIM is a data-driven pseudo-simulation framework and benchmark for autonomous vehicle planning, evaluating driving agents non-reactively… | 50 | 1076 | active |
| jeffbass/imagezmq imageZMQ is a set of Python classes that transport OpenCV images between computers using PyZMQ messaging. It enables distributed computer v… | 62 | 1069 | stable |
| InternRobotics/InternNav InternNav is an open-source PyTorch-based toolbox for building embodied navigation foundation models, supporting vision-language navigation… | 58 | 1061 | active |
| brenpoly/be-more-agent An offline-first conversational AI agent that turns a Raspberry Pi into a fully local voice assistant. It combines OpenWakeWord, Whisper.cp… | 64 | 1060 | active |
| open-gigaai/giga-models GigaModels is an open-source Python framework providing pipelines for training, inference, deployment, and compression of multi-modal, gene… | 62 | 1057 | active |
| InternRobotics/PointLLM PointLLM is a multimodal large language model that understands colored 3D point clouds of objects, built on a point cloud encoder fused wit… | 65 | 1051 | active |
| liuyuan-pal/SyncDreamer SyncDreamer is a synchronized multiview diffusion model that generates multiview-consistent images from a single-view image, released with … | 50 | 1045 | active |
| Xiaoqi-Zhao-DLUT/MSNet-M2SNet Official PyTorch implementations of MSNet and M2SNet, multi-scale subtraction networks for medical image segmentation such as polyp, lung i… | 74 | 1042 | active |
| videoflow/videoflow Videoflow is a Python framework for building distributed video and stream processing pipelines as directed acyclic graphs of producers, pro… | 66 | 1034 | active |
| podgorskiy/ALAE Official PyTorch implementation of Adversarial Latent Autoencoders (ALAE/StyleALAE), a CVPR 2020 paper combining autoencoders with GAN trai… | 32 | 3511 | maintenance |
| tensorlayer/SRGAN Reference implementation of SRGAN, a generative adversarial network for photo-realistic single image super-resolution, built on TensorLayer… | 23 | 3466 | maintenance |
| abizovnuralem/go2_ros2_sdk An unofficial ROS2 SDK for the Unitree Go2 quadruped robot (AIR/PRO/EDU), connecting over WebRTC (Wi-Fi) or CycloneDDS (Ethernet). It provi… | 57 | 1020 | active |
| bowang-lab/U-Mamba U-Mamba is a hybrid CNN-state-space-model (Mamba) network for biomedical image segmentation, built on top of the nnU-Net framework. It comb… | 26 | 1019 | active |
| shaoanlu/faceswap-GAN A Jupyter Notebook-based implementation of face swapping using a denoising autoencoder architecture enhanced with adversarial losses, VGGFa… | 32 | 3416 | maintenance |
| meetps/pytorch-semseg A PyTorch library implementing popular semantic segmentation architectures such as FCN, U-Net, SegNet, PSPNet, ICNet, FRRN, and LinkNet, wi… | 23 | 3402 | maintenance |
| huawei-noah/noah-research A collection of research code subprojects released by Huawei Noah's Ark Lab, each in its own directory. It is not an official Huawei produc… | 76 | 1004 | active |
| siyuanchen0214/Scam-AI-Multi-modal-Evaluation-System A Python-based multi-modal AI system for detecting fraudulent content across text, image, audio, and video, with provenance tracing and cro… | 39 | 1001 | active |
| catalyst-team/catalyst Catalyst is a high-level PyTorch framework for deep learning research and development, focused on reproducibility, rapid experimentation, a… | 64 | 3382 | maintenance |
| aserbao/AndroidCamera An Android library and demo app implementing a TikTok-style custom camera with video and audio editing features such as segment recording, … | 23 | 3297 | maintenance |
| zhanghang1989/ResNeSt ResNeSt is a PyTorch implementation of the Split-Attention Network, a ResNet variant that applies channel-wise attention across network bra… | 23 | 3261 | maintenance |
| jacksonliam/mjpg-streamer A lightweight command-line streaming tool that copies JPEG frames from input plugins (webcams, Raspberry Pi camera, files, OpenCV) to outpu… | 32 | 3247 | maintenance |
| Tramac/awesome-semantic-segmentation-pytorch A PyTorch library providing concise, modifiable reference implementations of many semantic segmentation models such as FCN, PSPNet, DeepLab… | 32 | 3069 | maintenance |
| cvlab-columbia/zero123 Zero-1-to-3 is a research codebase and pretrained diffusion model from Columbia CVLab that changes the camera viewpoint of an object from a… | 30 | 3058 | maintenance |
| Kurento/kurento-media-server Kurento Media Server is a C++/GStreamer-based media server that handles media transmission, processing, recording, and streaming over WebRT… | 10 | 3053 | maintenance |
| jfzhang95/pytorch-deeplab-xception A PyTorch implementation of the DeepLab v3+ semantic segmentation model with support for multiple backbones (Xception, ResNet, MobileNet, D… | 32 | 3000 | maintenance |
| tensorflow/graphics TensorFlow Graphics is a library of differentiable graphics layers for TensorFlow, including differentiable renderers, spatial transformers… | 64 | 2781 | maintenance |
| VainF/DeepLabV3Plus-Pytorch A PyTorch library providing pretrained DeepLabv3 and DeepLabv3+ semantic segmentation models for Pascal VOC and Cityscapes datasets. It inc… | 32 | 2699 | maintenance |
| yerfor/GeneFace GeneFace is the official PyTorch implementation of an ICLR 2023 paper on generalized, high-fidelity audio-driven 3D talking face synthesis … | 21 | 2657 | maintenance |
| HypoX64/DeepMosaics DeepMosaics is a Python application that automatically removes or adds mosaics in images and videos using semantic segmentation and image-t… | 23 | 2633 | maintenance |
| ShawnBIT/UNet-family A curated collection of UNet-family semantic segmentation models with PyTorch implementations and links to original papers and third-party … | 32 | 2592 | maintenance |
| CASIA-IVA-Lab/DANet DANet is the official PyTorch implementation of 'Dual Attention Network for Scene Segmentation' (CVPR 2019), which uses position and channe… | 32 | 2463 | maintenance |
| zai-org/CogVLM2 CogVLM2 is an open-source multi-modal vision-language model family built on Meta-Llama-3-8B-Instruct, offering image and video understandin… | 28 | 2433 | maintenance |
| iPERDance/iPERCore Impersonator++ (iPERCore) is a PyTorch implementation of Liquid Warping GAN with Attention, a unified framework for human image synthesis. … | 32 | 2393 | maintenance |
| nbei/Deep-Flow-Guided-Video-Inpainting A PyTorch implementation of the CVPR 2019 paper 'Deep Flow-Guided Video Inpainting', which fills missing regions in videos by completing op… | 32 | 2375 | maintenance |
| michuanhaohao/reid-strong-baseline A PyTorch implementation of the 'Bag of Tricks and A Strong Baseline for Deep Person Re-identification' paper (CVPRW 2019), providing end-t… | 32 | 2355 | maintenance |
| hzwer/ICCV2019-LearningToPaint A PyTorch research implementation of the ICCV 2019 paper 'Learning to Paint With Model-based Deep Reinforcement Learning'. It trains agents… | 41 | 2305 | maintenance |
| donnyyou/torchcv TorchCV is a PyTorch-based framework providing reimplementations of deep learning models for major computer vision tasks. It covers image c… | 32 | 2251 | maintenance |
| bigmb/Unet-Segmentation-Pytorch-Nest-of-Unets A PyTorch implementation of several U-Net variants for image segmentation, including UNet, R2U-Net, Attention U-Net, Attention R2U-Net, and… | 32 | 2249 | maintenance |
| idealo/image-quality-assessment A Python implementation of Google's NIMA (Neural Image Assessment) models that predict the aesthetic and technical quality of images using … | 10 | 2243 | maintenance |
| atriumlts/subpixel A TensorFlow reimplementation of the efficient sub-pixel convolutional neural network (ESPCN) for single-image super-resolution, based on S… | 32 | 2123 | maintenance |
| qubvel/efficientnet A Keras and TensorFlow Keras reimplementation of the EfficientNet convolutional neural network family (B0-B7), including ImageNet-pretraine… | 23 | 2100 | maintenance |
| ozan-oktay/Attention-Gated-Networks A PyTorch implementation of attention gates for convolutional neural networks, applied to U-Net and VGG-16 architectures. It targets medica… | 32 | 2064 | maintenance |
| apple/ml-fastvit Official PyTorch implementation of FastViT, a fast hybrid vision transformer architecture using structural reparameterization, published at… | 28 | 2027 | maintenance |
| JDAI-CV/FaceX-Zoo FaceX-Zoo is a PyTorch toolbox for face recognition that provides training modules with various state-of-the-art supervisory heads and back… | 32 | 1999 | maintenance |
| albertpumarola/GANimation Official PyTorch implementation of GANimation, an ECCV'18 research paper that animates facial expressions in a single image using a GAN con… | 32 | 1985 | maintenance |
| apple/ml-cvnets CVNets is Apple's open-source PyTorch library for training computer vision networks, covering classification, detection, segmentation, vide… | 32 | 1983 | maintenance |
| WuJie1010/Facial-Expression-Recognition.Pytorch A PyTorch implementation of CNN-based facial expression recognition achieving state-of-the-art accuracy on FER2013 (73.112%) and CK+ (94.64… | 32 | 1976 | maintenance |
| SummitKwan/transparent_latent_gan TL-GAN is a Python/TensorFlow project that makes a GAN's latent space transparent by discovering feature axes, enabling controlled image sy… | 32 | 1973 | maintenance |
| ronghuaiyang/arcface-pytorch A PyTorch implementation of ArcFace, a deep metric learning approach for face recognition that adds angular margin penalties to face embedd… | 32 | 1901 | maintenance |
| TreB1eN/InsightFace_Pytorch A PyTorch reimplementation of InsightFace/ArcFace for face recognition, including backbone models (IR-SE50, MobileFacenet) and pretrained w… | 32 | 1894 | maintenance |
| open-mmlab/mmaction MMAction is an open-source PyTorch toolbox for video action understanding, covering action recognition, temporal action detection, and spat… | 32 | 1876 | maintenance |
| NVIDIA/semantic-segmentation NVIDIA's PyTorch monorepo implementing the paper 'Hierarchical Multi-Scale Attention for Semantic Segmentation', with pretrained weights an… | 32 | 1828 | maintenance |
| yassouali/pytorch-segmentation A PyTorch library implementing multiple semantic segmentation models (DeepLab V3+, PSPNet, U-Net, SegNet, FCN, ENet, and others) with datas… | 26 | 1818 | maintenance |
| MCG-NJU/VideoMAE Official PyTorch implementation of VideoMAE, a masked autoencoder method for data-efficient self-supervised video pre-training with video t… | 32 | 1784 | maintenance |
| cszn/DnCNN DnCNN is the official implementation of the TIP 2017 paper 'Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising',… | 32 | 1727 | maintenance |
| bubbliiiing/unet-pytorch A PyTorch implementation of the U-Net semantic segmentation model with training, prediction, and mIoU evaluation scripts. It supports multi… | 23 | 1725 | maintenance |
| svip-lab/impersonator A PyTorch implementation of Liquid Gating GAN (ICCV 2019) that performs human motion imitation, appearance transfer, and novel view synthes… | 32 | 1717 | maintenance |
| dunbar12138/pix2pix3D pix2pix3D is the official PyTorch implementation of a CVPR 2023 paper on 3D-aware conditional image synthesis. It generates 3D objects (neu… | 31 | 1716 | maintenance |
| Temporal Segment Networks (TSN) Official code and pretrained models for Temporal Segment Networks (TSN), a deep learning framework for video action recognition published a… | 32 | 1577 | maintenance |
| HumanAIGC/EMO EMO (Emote Portrait Alive) is a research codebase from Alibaba's Institute for Intelligent Computing that generates expressive talking port… | 25 | 7594 | experimental |
| sniklaus/3d-ken-burns A PyTorch reference implementation of the 3D Ken Burns Effect from a Single Image paper, which animates a still photo with a virtual camera… | 70 | 1569 | maintenance |
| AlexHex7/Non-local_pytorch A PyTorch implementation of the Non-local Neural Block from the paper 'Non-local Neural Networks', providing multiple variants (concatenati… | 32 | 1564 | maintenance |
| open-mmlab/Multimodal-GPT Multimodal-GPT is an open-source project for training a multimodal chatbot that combines vision and language instructions, built on OpenFla… | 30 | 1512 | maintenance |
| szagoruyko/attention-transfer PyTorch reference implementation of the ICLR 2017 paper 'Paying More Attention to Attention', which improves convolutional neural networks … | 32 | 1463 | maintenance |
| datitran/face2face-demo A pix2pix demo that learns from facial landmarks and translates them into a trained target face, with a real-time webcam application. It in… | 32 | 1461 | maintenance |
| facebookresearch/MaskFormer MaskFormer is a PyTorch/Detectron2-based implementation of the NeurIPS 2021 paper 'Per-Pixel Classification is Not All You Need for Semanti… | 10 | 1460 | maintenance |
| hysts/pytorch_image_classification A PyTorch library implementing many image classification architectures (ResNet, DenseNet, WRN, PyramidNet, SENet, etc.) and augmentation te… | 10 | 1449 | maintenance |
| megvii-research/ML-GCN A PyTorch implementation of ML-GCN, the CVPR 2019 paper 'Multi-Label Image Recognition with Graph Convolutional Networks'. It provides trai… | 32 | 1446 | maintenance |
| BachiLi/redner redner is a differentiable Monte Carlo ray tracer that computes exact gradients of rendered images with respect to arbitrary scene paramete… | 32 | 1444 | maintenance |
| xuebinqin/BASNet BASNet is the official PyTorch implementation of the CVPR 2019 paper 'BASNet: Boundary-Aware Salient Object Detection', a deep learning mod… | 32 | 1436 | maintenance |
| vlfeat/matconvnet MatConvNet is a MATLAB toolbox implementing convolutional neural networks (CNNs) for computer vision applications. It supports training and… | 32 | 1430 | maintenance |
| rmokady/CLIP_prefix_caption Official implementation of ClipCap, a CLIP-based image captioning model that maps CLIP image encodings to a GPT-2 prefix to generate captio… | 32 | 1423 | maintenance |
| gradslam/gradslam gradslam is a fully differentiable dense SLAM library built on PyTorch, providing differentiable building blocks such as nonlinear least sq… | 23 | 1422 | maintenance |
| yu-changqian/TorchSeg A fast, modular PyTorch reference implementation for training and evaluating semantic segmentation models such as FCN, DFN, BiSeNet, PSPNet… | 23 | 1410 | maintenance |
| open-mmlab/mmfashion MMFashion is an open-source PyTorch-based toolbox for visual fashion analysis from the OpenMMLab project. It provides modular implementatio… | 32 | 1369 | maintenance |
| orobix/retina-unet A Python implementation of a U-Net convolutional neural network for segmenting blood vessels in retina fundus images. It performs binary pi… | 32 | 1354 | maintenance |
| aitorzip/PyTorch-CycleGAN A clean, readable PyTorch implementation of CycleGAN for unpaired image-to-image translation using cycle-consistent adversarial networks. I… | 32 | 1319 | maintenance |
| hukkelas/DeepPrivacy DeepPrivacy is a PyTorch-based GAN that automatically anonymizes faces in images and videos by generating realistic synthetic replacements.… | 32 | 1316 | maintenance |
| d-li14/involution Official PyTorch implementation of the involution neural operator from the CVPR 2021 paper 'Involution: Inverting the Inherence of Convolut… | 32 | 1310 | maintenance |
| bubbliiiing/deeplabv3-plus-pytorch A PyTorch implementation of the DeepLabv3+ semantic segmentation model with MobileNetV2 and Xception backbones. It includes scripts for tra… | 23 | 1292 | maintenance |
| YuanxunLu/LiveSpeechPortraits A PyTorch implementation of the SIGGRAPH Asia 2021 paper 'Live Speech Portraits', which generates photorealistic personalized talking-head … | 32 | 1283 | maintenance |
| elanmart/cbp-translate A demo application that live-translates foreign-language speech in videos into subtitles, mimicking the Cyberpunk 2077 translation effect. … | 32 | 1274 | maintenance |
| marcbelmont/cnn-watermark-removal A TensorFlow implementation of a fully convolutional neural network that removes transparent watermark overlays from images. It trains on s… | 32 | 1269 | maintenance |
| DrSleep/tensorflow-deeplab-resnet A TensorFlow re-implementation of the DeepLab-ResNet model for semantic image segmentation, trained and evaluated on the PASCAL VOC dataset… | 10 | 1258 | maintenance |