function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| alibaba/Tora Tora is Alibaba's official implementation of a trajectory-oriented Diffusion Transformer (DiT) for controllable video generation, integrati… | 64 | 1241 | active |
| Roblox/cube Cube is Roblox's open-source family of foundation models for 3D intelligence, including text-to-3D shape generation and part-controllable m… | 58 | 1241 | active |
| XPixelGroup/HYPIR Official PyTorch implementation of HYPIR, a SIGGRAPH 2025 method that harnesses diffusion-yielded score priors for image restoration. It pr… | 39 | 1241 | active |
| jcjohnson/fast-neural-style A Torch (Lua) implementation of feedforward neural style transfer from the ECCV 2016 paper 'Perceptual Losses for Real-Time Style Transfer … | 32 | 4359 | maintenance |
| facebookresearch/deit Official PyTorch repository for DeiT and related vision transformer architectures (CaiT, ResMLP, PatchConvnet, DeiT III), providing trainin… | 10 | 4355 | maintenance |
| declare-lab/tango Tango is a family of latent diffusion models for text-to-audio generation, with Tango 2 improving prompt alignment via DPO-based fine-tunin… | 45 | 1239 | active |
| apchenstu/TensoRF TensoRF is a PyTorch implementation of the ECCV 2022 paper 'TensoRF: Tensorial Radiance Fields', which models and reconstructs radiance fie… | 44 | 1239 | stable |
| X-Square-Robot/wall-x Wall-X is the open-source training and inference stack for X Square Robot's WALL series of embodied foundation models (VLAs) for general-pu… | 61 | 1236 | active |
| meta-pytorch/attention-gym Attention Gym is a collection of tools, examples, and reference implementations for working with PyTorch's FlexAttention API. It provides a… | 91 | 1234 | active |
| chengzeyi/Comfy-WaveSpeed A ComfyUI custom node plugin that acts as an all-in-one inference optimization solution for diffusion models, built around First Block Cach… | 66 | 1231 | stable |
| lucidrains/deep-daze Deep Daze is a simple command line tool for text-to-image generation that combines OpenAI's CLIP with a Siren implicit neural representatio… | 23 | 4315 | maintenance |
| MoonshotAI/FlashKDA FlashKDA is a set of high-performance CUDA kernels (built on CUTLASS) implementing Kimi Delta Attention, a linear attention mechanism, for … | 57 | 1229 | active |
| CSSLab/maia-chess Maia is a collection of human-like neural network chess engines trained on millions of human games, targeting skill levels from ELO 1100 to… | 61 | 1228 | active |
| Tencent-Hunyuan/HunyuanCustom HunyuanCustom is a multimodal-driven customized video generation framework built on HunyuanVideo, supporting image, text, audio, and video … | 40 | 1227 | active |
| NVIDIA/BigVGAN BigVGAN is NVIDIA's official PyTorch implementation of a universal neural vocoder (ICLR 2023) that generates high-fidelity raw audio wavefo… | 23 | 1227 | stable |
| alibaba/x-deeplearning X-DeepLearning (XDL) is an industrial deep learning framework from Alibaba optimized for high-dimension sparse data scenarios such as adver… | 23 | 4304 | maintenance |
| MoonshotAI/Kimi-VL Kimi-VL is an open-source Mixture-of-Experts vision-language model (VLM) with a 2.8B activated parameter language decoder, offering multimo… | 33 | 1224 | active |
| ElectricAlexis/NotaGen NotaGen is a symbolic music generation model that produces high-quality classical sheet music using LLM-style training paradigms: pre-train… | 32 | 1223 | active |
| eduardoleao052/js-pytorch JS-PyTorch is a deep learning library for JavaScript that closely mirrors PyTorch's syntax, providing tensor operations, automatic differen… | 16 | 1222 | active |
| SciML/NeuralPDE.jl NeuralPDE.jl is a Julia library of physics-informed neural network (PINN) solvers for ordinary, stochastic, and partial differential equati… | 99 | 1220 | active |
| MotrixLab/SMPLer-X Official code for SMPLer-X, a family of foundation models for expressive human pose and shape estimation (EHPS) that unifies body, hand, an… | 59 | 1220 | stable |
| Project-MONAI/research-contributions A collection of peer-reviewed research prototype implementations built on the MONAI framework for medical imaging AI. It serves as a fast-t… | 43 | 1220 | active |
| zkonduit/ezkl EZKL is a Rust-based library and command-line tool that converts deep learning models and arbitrary computational graphs (exported as ONNX)… | 75 | 1219 | active |
| Aratako/Irodori-TTS Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer… | 58 | 1219 | active |
| lucidrains/perceiver-pytorch A PyTorch implementation of the Perceiver architecture (General Perception with Iterative Attention) and its follow-up Perceiver IO. It pro… | 62 | 1217 | active |
| fudan-generative-vision/champ Champ is a research framework for controllable and consistent human image animation using 3D parametric guidance (SMPL-based depth, normal,… | 25 | 4261 | maintenance |
| Picsart-AI-Research/Text2Video-Zero Official implementation of Text2Video-Zero, a zero-shot text-to-video generation method that adapts text-to-image diffusion models like Sta… | 30 | 4245 | maintenance |
| ifzhang/FairMOT FairMOT is a research implementation of a one-shot multi-object tracking model that jointly performs object detection and re-identification… | 32 | 4244 | maintenance |
| hasktorch/hasktorch Hasktorch is a Haskell library for tensor math and neural networks, built on bindings to the C++ libtorch libraries that power PyTorch. It … | 76 | 1211 | active |
| datawhalechina/torch-rechub Torch-RecHub is a lightweight PyTorch framework for building recommendation system models with 30+ out-of-the-box algorithms covering ranki… | 94 | 1207 | active |
| DachunKai/EvTexture Official PyTorch implementation of EvTexture and EvTexture++, event-driven video super-resolution models that use event-camera signals to e… | 54 | 1207 | active |
| willisma/SiT Official PyTorch implementation of Scalable Interpolant Transformers (SiT), a family of generative models built on Diffusion Transformers t… | 53 | 1206 | active |
| metavoiceio/metavoice-src MetaVoice-1B is a 1.2B parameter foundational text-to-speech model trained on 100K hours of speech, focused on emotional rhythm and tone in… | 26 | 4205 | maintenance |
| Artelnics/opennn OpenNN is an open-source C++ library for building, training, and deploying neural networks for advanced analytics. It is dependency-free, o… | 97 | 1198 | active |
| openvinotoolkit/nncf NNCF is Intel's Neural Network Compression Framework, a Python library providing post-training and training-time compression algorithms (qu… | 94 | 1197 | active |
| Calamari-OCR/calamari Calamari is a Python-based OCR engine for line-based automatic text recognition, built on OCRopy and Kraken with a TensorFlow deep-learning… | 74 | 1197 | active |
| autonomousvision/stylegan-t Official training code for StyleGAN-T, an ICML 2023 paper on fast large-scale text-to-image synthesis using GANs. It provides dataset prepa… | 31 | 1197 | active |
| DeepRec-AI/DeepRec DeepRec is a high-performance deep learning framework for recommendation models, built on TensorFlow 1.15 with Intel and NVIDIA TensorFlow … | 24 | 1197 | active |
| martinpacesa/BindCraft BindCraft is a Python-based computational pipeline for de novo protein binder design that combines AlphaFold2 backpropagation, ProteinMPNN,… | 78 | 1196 | active |
| haidog-yaqub/MeanFlow An unofficial PyTorch implementation of MeanFlow and iMF, one-step generative modeling methods based on flow matching. It provides config-d… | 59 | 1196 | active |
| deepseek-ai/DeepSeek-VL DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf… | 25 | 4175 | maintenance |
| EvolvingLMMs-Lab/LLaVA-OneVision-2 A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis… | 72 | 1195 | active |
| facebookresearch/esm Meta FAIR's Evolutionary Scale Modeling (ESM) library providing Transformer protein language models with pretrained weights, including ESM-… | 10 | 4170 | maintenance |
| Cerebras/modelzoo Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and … | 77 | 1193 | active |
| Tencent-Hunyuan/HunyuanWorld-Mirror HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds… | 54 | 1191 | active |
| shallowdream204/DreamClear DreamClear is a diffusion-transformer based real-world image restoration model for high-fidelity super-resolution, published at NeurIPS 202… | 27 | 1191 | active |
| a-r-j/graphein Graphein is a Python library for constructing graph and mesh representations of proteins, RNA, molecules, and biological interaction networ… | 76 | 1190 | active |
| ali-vilab/UniAnimate UniAnimate is the official code for a research paper on animating a reference human image into a video that follows a driving pose sequence… | 31 | 1189 | active |
| SHI-Labs/Neighborhood-Attention-Transformer Official PyTorch implementation of the Neighborhood Attention Transformer (NAT/DiNAT), a family of efficient vision transformers with local… | 32 | 1184 | stable |
| mlmed/torchxrayvision TorchXRayVision is an open-source PyTorch library providing pre-trained deep learning models and a unified interface for publicly available… | 91 | 1183 | active |
| chensjtu/GaussianObject GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,… | 26 | 1183 | active |
| msracver/Deformable-ConvNets Official MXNet implementation of Deformable Convolutional Networks (ICCV 2017) and R-FCN, including deformable convolution and ROI pooling … | 32 | 4121 | maintenance |
| XPixelGroup/DiffBIR DiffBIR is a blind image restoration framework that uses generative diffusion priors to restore degraded real-world images. It provides pre… | 34 | 4119 | maintenance |
| mlfoundations/open_flamingo OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima… | 23 | 4118 | maintenance |
| DAMO-NLP-SG/VideoLLaMA3 VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de… | 37 | 1179 | active |
| ubicomplab/rPPG-Toolbox rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and… | 51 | 1178 | active |
| NVlabs/Deep_Object_Pose NVIDIA's Deep Object Pose Estimation (DOPE), a deep learning system for detecting known objects and estimating their 6-DoF pose from RGB ca… | 48 | 1178 | active |
| balancap/SSD-Tensorflow A TensorFlow re-implementation of the Single Shot MultiBox Detector (SSD) for object detection, including VGG-based SSD-300 and SSD-512 net… | 32 | 4101 | maintenance |
| AlgRUC/JittorGeometric JittorGeometric is a graph machine learning library built on the Jittor deep learning framework, providing implementations of 40+ Graph Neu… | 62 | 1177 | active |
| tjiiv-cprg/EPro-PnP EPro-PnP is a probabilistic Perspective-n-Points (PnP) layer for end-to-end 6DoF monocular object pose estimation networks, built on PyTorc… | 41 | 1175 | stable |
| princeton-nlp/MeZO MeZO is a memory-efficient zeroth-order optimizer that fine-tunes language models using only forward passes, with the same memory footprint… | 29 | 1173 | stable |
| bytedance/1d-tokenizer A research repository from ByteDance containing code and pretrained model weights for 1D visual tokenizers (TiTok, TA-TiTok, FlowTok) and i… | 29 | 1172 | active |
| csguoh/MambaIR MambaIR and MambaIRv2 are PyTorch-based image restoration models built on Mamba state-space models, published at ECCV 2024 and CVPR 2025. T… | 54 | 1171 | active |
| magicleap/SuperGluePretrainedNetwork SuperGlue is a PyTorch implementation of a graph neural network with an optimal matching layer that matches sparse image features between t… | 32 | 4072 | maintenance |
| baidu-research/warp-ctc A fast parallel implementation of the Connectionist Temporal Classification (CTC) loss function for CPU and CUDA GPU, with a simple C inter… | 32 | 4069 | maintenance |
| iver56/torch-audiomentations A PyTorch library for fast audio data augmentation, inspired by audiomentations. It provides GPU-accelerated, differentiable audio transfor… | 58 | 1167 | active |
| SuperBruceJia/EEG-DL EEG-DL is a deep learning library built on TensorFlow for classifying EEG signals, supporting many architectures including CNNs, RNNs, GCNs… | 47 | 1167 | active |
| nv-tlabs/LLaMA-Mesh LLaMA-Mesh is a fine-tuned large language model from NVIDIA Research that generates and understands 3D meshes by representing vertex coordi… | 28 | 1166 | active |
| sksq96/pytorch-summary A PyTorch library providing a Keras-style model.summary() that prints layer types, output shapes, parameter counts, and memory estimates. I… | 32 | 4053 | maintenance |
| facebookresearch/VideoPose3D A PyTorch implementation of CVPR 2019 research on 3D human pose estimation in video using temporal convolutions over 2D keypoint trajectori… | 10 | 4052 | maintenance |
| yosinski/deep-visualization-toolbox A GUI toolbox for visualizing and understanding deep neural networks, showing per-unit activations, backprop/deconv, and regularized-optimi… | 32 | 4051 | maintenance |
| Text-to-Audio/AudioLCM AudioLCM is a PyTorch implementation of an ACM-MM'24 paper for efficient, high-quality text-to-audio generation using latent consistency mo… | 37 | 1165 | active |
| thunlp/OpenKE OpenKE is an open-source PyTorch-based toolkit for knowledge graph embedding (knowledge representation learning), with C++ accelerated data… | 32 | 4047 | maintenance |
| majianjia/nnom NNoM is a high-level neural network inference library written in C for microcontrollers. It converts Keras models into optimized on-device … | 23 | 1164 | stable |
| CyberAgentAILab/TANGO TANGO is a research library from CyberAgent AI Lab that generates co-speech gesture videos by reenactment, using hierarchical audio-motion … | 39 | 1163 | active |
| facebookresearch/encodec EnCodec is a deep learning based neural audio codec from Meta AI that compresses mono 24 kHz and stereo 48 kHz audio to bitrates from 1.5 t… | 32 | 4041 | maintenance |
| 3D ResNets for Action Recognition A PyTorch implementation of 3D ResNet and R(2+1)D models for video action recognition, accompanying CVPR 2018 and related papers. It includ… | 23 | 4038 | maintenance |
| MCG-NKU/E2FGVI E2FGVI is the official PyTorch implementation of the CVPR 2022 paper 'Towards An End-to-End Framework for Flow-Guided Video Inpainting'. It… | 32 | 1161 | stable |
| minivision-ai/photo2cartoon A Python deep-learning project from Minivision that converts real portrait photos into cartoon-style avatars using unpaired image translati… | 32 | 4029 | maintenance |
| LibCity/Bigscity-LibCity LibCity is an open-source PyTorch library for urban spatial-temporal data mining, providing a unified pipeline for traffic prediction resea… | 23 | 1157 | active |
| fundamentalvision/Deformable-DETR Official PyTorch implementation of Deformable DETR, an efficient end-to-end object detector that uses deformable attention to fix DETR's sl… | 32 | 4015 | maintenance |
| sirius-ai/LPRNet_Pytorch A PyTorch implementation of LPRNet, a lightweight deep neural network for license plate recognition. It ships with pretrained weights focus… | 32 | 1156 | stable |
| gemelo-ai/vocos Vocos is a fast neural vocoder that synthesizes audio waveforms from acoustic features such as mel-spectrograms or EnCodec tokens. It uses … | 65 | 1155 | stable |
| deepmodeling/Uni-Mol Uni-Mol is a collection of 3D molecular representation learning frameworks and pretrained models for tasks like molecule property predictio… | 34 | 1155 | active |
| JunMa11/SegLossOdyssey A curated collection of loss functions for medical image segmentation, accompanying the 'Loss Odyssey in Medical Image Segmentation' survey… | 32 | 4007 | maintenance |
| SystemErrorWang/White-box-Cartoonization Official TensorFlow implementation of the CVPR 2020 paper 'Learning to Cartoonize Using White-box Cartoon Representations', which converts … | 61 | 4001 | maintenance |
| TJU-Aerial-Robotics/YOPO YOPO is a learning-based one-stage planner for quadrotor autonomous navigation in obstacle-dense environments, integrating perception, mapp… | 79 | 1153 | active |
| OpenGVLab/VisionLLM VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc… | 33 | 1153 | active |
| quark0/darts DARTS is the official PyTorch implementation of the ICLR 2019 paper 'DARTS: Differentiable Architecture Search', which performs neural arch… | 32 | 3997 | maintenance |
| TensorSpeech/TensorFlowTTS TensorFlowTTS is a Python library providing real-time state-of-the-art text-to-speech architectures (Tacotron-2, FastSpeech/FastSpeech2, Me… | 23 | 3995 | maintenance |
| amazon-science/mm-cot Official PyTorch implementation of the paper 'Multimodal Chain-of-Thought Reasoning in Language Models', which adds vision features to a tw… | 31 | 3985 | maintenance |
| JDAI-CV/fast-reid FastReID is a PyTorch-based research platform implementing state-of-the-art re-identification algorithms for persons, vehicles, and faces. … | 23 | 3981 | maintenance |
| open-gigaai/giga-world-1 GigaWorld-1 is an open-source framework providing training, inference, data processing, checkpoint conversion, and LoRA merge workflows for… | 54 | 1147 | active |
| ShiqiYu/OpenGait OpenGait is a flexible and extensible Python framework for gait recognition research, providing implementations of state-of-the-art models … | 67 | 1146 | active |
| ScorpioLea/AiCE AiCE is a Python tool that predicts high-fitness protein mutations by sampling sequences from protein inverse folding models such as Protei… | 40 | 1144 | active |
| HengyiWang/spann3r Spann3R is a transformer-based model for dense 3D reconstruction from ordered or unordered image collections, built on the DUSt3R paradigm.… | 26 | 1141 | active |
| cvg/glue-factory Glue Factory is a PyTorch-based library for training and evaluating deep neural networks that detect and match local visual features (point… | 69 | 1140 | active |
| THUMNLab/AutoGL AutoGL is an autoML framework and toolkit for machine learning on graphs, built on PyTorch with PyTorch Geometric and DGL backends. It prov… | 47 | 1140 | active |
| rohitgandikota/sliders Official implementation of Concept Sliders, LoRA adaptors that enable precise, plug-and-play control of attributes in diffusion models like… | 52 | 1139 | active |
| clovaai/deep-text-recognition-benchmark Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio… | 32 | 3942 | maintenance |