function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| lyhue1991/torchkeras torchkeras is a lightweight PyTorch model training template library that brings Keras-style compile/fit/evaluate APIs to PyTorch. Its core … | 55 | 2009 | active |
| AntixK/PyTorch-VAE A collection of Variational Autoencoder (VAE) model implementations in PyTorch, including Beta-VAE, VQ-VAE, IWAE, WAE, and others, with a f… | 38 | 7665 | maintenance |
| kohya-ss/musubi-tuner Musubi Tuner is a set of Python scripts for training LoRA (Low-Rank Adaptation) adapters for video and image generation model architectures… | 85 | 2002 | active |
| bytetriper/RAE Official PyTorch implementation of 'Diffusion Transformers with Representation Autoencoders' (RAE), a two-stage image generation pipeline u… | 48 | 2001 | active |
| mil-tokyo/webdnn WebDNN is a framework for running deep neural network inference directly in the web browser, accepting ONNX models without Python preproces… | 65 | 1999 | active |
| clementchadebec/benchmark_VAE Pythae is a PyTorch library that unifies implementations of many Variational Autoencoder (VAE) variants under a common interface, enabling … | 23 | 1995 | active |
| facebookresearch/dino PyTorch implementation of DINO, a self-supervised learning method for training Vision Transformers, with pretrained model weights. It is th… | 10 | 7611 | maintenance |
| chaidiscovery/chai-lab Chai-1 is a state-of-the-art multi-modal foundation model for biomolecular structure prediction, handling proteins, small molecules, DNA, R… | 65 | 1986 | active |
| kwuking/TimeMixer Official PyTorch implementation of TimeMixer, an ICLR 2024 model for time series forecasting using decomposable multiscale mixing. It has s… | 46 | 1981 | active |
| lucidrains/titans-pytorch An unofficial PyTorch implementation of the Titans architecture, a neural long-term memory module for transformers that learns to memorize … | 70 | 1980 | active |
| davda54/sam An unofficial PyTorch implementation of Sharpness-Aware Minimization (SAM) and its adaptive variant ASAM, provided as an optimizer wrapper … | 32 | 1980 | stable |
| cloneofsimo/lora A Python library for applying Low-Rank Adaptation (LoRA) to quickly fine-tune text-to-image diffusion models like Stable Diffusion. It prod… | 22 | 7550 | maintenance |
| adobe-research/custom-diffusion Custom Diffusion is a research codebase for efficiently fine-tuning text-to-image diffusion models like Stable Diffusion on a few example i… | 69 | 1978 | stable |
| JIA-Lab-research/DreamOmni2 DreamOmni2 is the official PyTorch implementation of a CVPR 2026 Highlight model for multimodal instruction-based image editing and generat… | 51 | 1978 | active |
| showlab/Show-o Show-o is a research repository implementing a unified transformer model that combines autoregressive and discrete diffusion modeling for m… | 50 | 1973 | active |
| WassimTenachi/PhySO PhySO is a Python library for physical symbolic optimization that uses deep reinforcement learning to discover analytical physical laws fro… | 52 | 1971 | active |
| google-deepmind/tapnet Google DeepMind's official repository for Tracking Any Point (TAP), containing the TAP-Vid and TAPVid-3D benchmarks, the TAPIR and TAPNext … | 74 | 1968 | active |
| meta-recsys/generative-recommenders Meta's research library implementing HSTU and M-FALCON from the ICML'24 paper 'Actions Speak Louder than Words: Trillion-Parameter Sequenti… | 69 | 1966 | active |
| OpenMotionLab/MotionGPT MotionGPT is a unified motion-language model that treats 3D human motion as a foreign language by converting motion into discrete motion to… | 32 | 1961 | active |
| 2U1/Qwen-VL-Series-Finetune An open-source Python repository providing training scripts for fine-tuning Alibaba's Qwen-VL series of vision-language models (Qwen2-VL, Q… | 67 | 1960 | active |
| open-mmlab/mmagic MMagic is OpenMMLab's toolbox for generative and multimodal AI image/video creation, built on PyTorch. It provides a large model zoo coveri… | 23 | 7457 | maintenance |
| SizheAn/PanoHead PanoHead is the official PyTorch implementation of a CVPR 2023 paper presenting a 3D-aware GAN that synthesizes geometry-aware, view-consis… | 29 | 1956 | active |
| alibaba/EasyCV EasyCV is an all-in-one PyTorch-based computer vision toolkit from Alibaba covering self-supervised learning, vision transformers, and majo… | 32 | 1954 | active |
| Fafa-DL/Awesome-Backbones A PyTorch-based framework that integrates many deep learning backbone models (CNNs and vision transformers like ResNet, EfficientNet, Swin … | 33 | 1953 | active |
| meta-pytorch/opacus Opacus is a PyTorch library for training neural networks with differential privacy via DP-SGD, requiring minimal code changes through its P… | 79 | 1952 | active |
| Yuliang-Liu/Monkey Monkey is a large multi-modal model (LMM) research project from CVPR 2024 that improves image understanding via higher input resolution and… | 65 | 1951 | active |
| LTH14/mar Official PyTorch implementation of MAR (Masked Autoregressive) image generation with DiffLoss, from the NeurIPS 2024 paper 'Autoregressive … | 54 | 1949 | stable |
| openai/guided-diffusion OpenAI's codebase for guided diffusion models from the paper 'Diffusion Models Beat GANs on Image Synthesis', including classifier conditio… | 32 | 7419 | maintenance |
| Vchitect/Latte Official PyTorch implementation of Latte, a latent diffusion transformer for video generation. It includes model definitions, pre-trained c… | 71 | 1948 | active |
| kyegomez/BitNet A PyTorch implementation of the BitNet architecture from the paper 'BitNet: Scaling 1-bit Transformers for Large Language Models', providin… | 72 | 1945 | active |
| Graph Convolutional Networks (GCN) A TensorFlow implementation of Graph Convolutional Networks (GCN) for semi-supervised node classification on graphs, accompanying the ICLR … | 32 | 7400 | maintenance |
| tensorlayer/TensorLayer TensorLayer is a TensorFlow-based deep learning and reinforcement learning library offering customizable neural layers for researchers and … | 23 | 7381 | maintenance |
| PixArt-alpha/PixArt-sigma PixArt-Σ is a PyTorch implementation of a diffusion transformer model for high-resolution (up to 4K) text-to-image generation, trained with… | 25 | 1939 | active |
| microsoft/Magma Magma is Microsoft Research's foundation model for multimodal AI agents, released as an 8B vision-language model that understands images an… | 53 | 1937 | active |
| NVlabs/RADIO Official PyTorch implementation of AM-RADIO and its successors (RADIOv2.5, C-RADIOv4), agglomerative vision foundation models distilled fro… | 64 | 1933 | active |
| tum-pbs/PhiFlow PhiFlow is an open-source Python simulation toolkit for solving partial differential equations with support for optimization and machine le… | 72 | 1929 | active |
| Audio-AGI/AudioSep AudioSep is the official implementation of the 'Separate Anything You Describe' foundation model for open-domain, language-queried audio so… | 28 | 1929 | active |
| Yuanshi9815/OminiControl OminiControl is a universal control framework for Diffusion Transformer models like FLUX, supporting subject-driven and spatial control (ed… | 62 | 1927 | active |
| OpenTalker/video-retalking VideoReTalking is a Python research system from SIGGRAPH Asia 2022 that edits real-world talking-head videos to match a given audio track, … | 23 | 7280 | maintenance |
| FACEGOOD/FACEGOOD-Audio2Face FACEGOOD Audio2Face is an open-source deep learning framework that converts audio into facial blendshape weights for driving digital humans… | 64 | 1909 | active |
| lucidrains/byol-pytorch A PyTorch library implementing the Bootstrap Your Own Latent (BYOL) self-supervised learning method from DeepMind. It wraps any image-based… | 58 | 1903 | active |
| google-deepmind/penzai Penzai is a JAX research toolkit for building, editing, and visualizing neural networks as legible, functional pytree data structures. It i… | 10 | 1901 | active |
| sapientinc/HRM-Text HRM-Text is a 1B-parameter text generation model based on the hierarchical recurrent HRM architecture, released with a complete pretraining… | 53 | 1899 | active |
| flexflow/flexflow-train FlexFlow Train is a deep learning framework that accelerates distributed DNN training by automatically searching for efficient parallelizat… | 67 | 1898 | active |
| showlab/ShowUI ShowUI is an open-source, lightweight 2B vision-language-action model for GUI agents and computer use, accepted at CVPR 2025. The repositor… | 57 | 1893 | active |
| zju3dv/GVHMR GVHMR is a research codebase implementing the SIGGRAPH Asia 2024 paper 'World-Grounded Human Motion Recovery via Gravity-View Coordinates'.… | 60 | 1882 | active |
| qqwweee/keras-yolo3 A Keras (TensorFlow backend) implementation of YOLOv3 for object detection, including Darknet weight conversion, image/video detection scri… | 32 | 7114 | maintenance |
| nitrain/nitrain Nitrain is a framework-agnostic Python library for sampling, augmenting, and training AI models on medical imaging datasets, with support f… | 23 | 1880 | active |
| NVIDIA-AI-IOT/Lidar_AI_Solution NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil… | 72 | 1867 | active |
| OpenNMT OpenNMT is an open-source ecosystem for neural machine translation and sequence learning, with PyTorch (OpenNMT-py) and TensorFlow (OpenNMT… | 44 | 7012 | maintenance |
| KlingAIResearch/ReCamMaster ReCamMaster is a reference implementation of a camera-controlled generative video rendering model that re-renders a single source video alo… | 44 | 1855 | active |
| facebookresearch/MetaCLIP Meta's research code and models for Meta CLIP, a reimplementation and scaling recipe for CLIP-style contrastive vision-language models, inc… | 82 | 1854 | active |
| SerpentAI/SerpentAI Serpent.AI is a Python framework for building game agents—AIs and bots that learn to play any video game you own—turning games into machine… | 10 | 6992 | maintenance |
| LuChengTHU/dpm-solver Official PyTorch implementation of DPM-Solver and DPM-Solver++, fast high-order ODE solvers for diffusion probabilistic model sampling that… | 32 | 1852 | stable |
| dotnet/TorchSharp TorchSharp is a .NET library providing bindings to LibTorch, the library that powers PyTorch, with a focus on tensors and a PyTorch-like AP… | 73 | 1850 | active |
| probcomp/Gen.jl Gen.jl is a general-purpose probabilistic programming system embedded in Julia that lets users write generative models as probabilistic pro… | 62 | 1850 | active |
| NVlabs/stylegan3 Official PyTorch implementation of StyleGAN3 (Alias-Free GANs), a state-of-the-art generative adversarial network for high-fidelity image s… | 32 | 6943 | maintenance |
| Tencent-Hunyuan/HunyuanVideo-I2V HunyuanVideo-I2V is Tencent's open-source image-to-video generation framework built on the HunyuanVideo diffusion model, providing PyTorch … | 54 | 1840 | active |
| ytongbai/LVM LVM is a large vision model trained with sequential next-token prediction over 'visual sentences', using no linguistic data. It builds on O… | 30 | 1838 | active |
| NVIDIA/pix2pixHD PyTorch implementation of pix2pixHD, a conditional GAN method for synthesizing and manipulating high-resolution (2048x1024) photorealistic … | 32 | 6923 | maintenance |
| clovaai/donut Donut is an OCR-free end-to-end Transformer model for visual document understanding tasks such as document classification and information e… | 23 | 6919 | maintenance |
| cazala/synaptic Synaptic is an architecture-free neural network library for JavaScript that runs in both Node.js and the browser. It supports building and … | 66 | 6912 | maintenance |
| dauparas/ProteinMPNN ProteinMPNN is a PyTorch-based tool that designs amino acid sequences for given protein backbone structures using a message-passing neural … | 23 | 1834 | stable |
| openai/point-e Point-E is OpenAI's official release of models and code for generating 3D point clouds from text prompts or images using diffusion models. … | 32 | 6895 | maintenance |
| timothybrooks/instruct-pix2pix PyTorch implementation of InstructPix2Pix, a diffusion-based model that edits images according to natural language instructions (e.g., 'tur… | 31 | 6885 | maintenance |
| boundless-large-model/boundless-world-model Boundless-World-Model (BWM) is a physically consistent, action-conditioned video world model built on Wan2.2-TI2V-5B that acts as a low-cos… | 59 | 1824 | active |
| Zheng-Chong/CatVTON CatVTON is a lightweight diffusion model for virtual try-on that swaps clothing onto a person image using a concatenation-based architectur… | 40 | 1824 | active |
| GestaltCogTeam/BasicTS BasicTS is a Python benchmark library and toolkit for fair and scalable time series analysis, built on PyTorch. It supports forecasting, cl… | 68 | 1814 | active |
| yerfor/GeneFacePlusPlus GeneFace++ is the official PyTorch implementation of a NeRF-based system for generalized and stable real-time 3D talking face generation. I… | 26 | 1809 | active |
| apple/ml-4m 4M is a framework from Apple and EPFL for training any-to-any multimodal foundation models using masked modeling over discrete tokens acros… | 35 | 1808 | active |
| Robbyant/lingbot-va LingBot-VA is an autoregressive diffusion framework that unifies video world modeling and robot policy learning in a single interleaved vid… | 56 | 1806 | active |
| microsoft/mattergen MatterGen is Microsoft's official implementation of a generative diffusion model for designing inorganic crystalline materials across the p… | 67 | 1801 | active |
| HuangJunJie2017/BEVDet BEVDet is a Python research codebase implementing the BEVDet series of bird's-eye-view (BEV) 3D object detection models for autonomous driv… | 23 | 1801 | active |
| zai-org/CogVLM CogVLM is an open-source visual language model (17B) combining a vision encoder with a pretrained language model for image understanding an… | 28 | 6744 | maintenance |
| QwenLM/Qwen-VL Official repository for Qwen-VL, Alibaba Cloud's large vision-language model family, including the pretrained Qwen-VL and instruction-tuned… | 28 | 6726 | maintenance |
| acids-ircam/RAVE RAVE is the official PyTorch implementation of a realtime audio variational autoencoder for fast, high-quality neural audio synthesis. It s… | 55 | 1790 | active |
| lehaifeng/T-GCN A collection of research source code implementing Temporal Graph Convolutional Networks (T-GCN) and related variants for urban traffic flow… | 50 | 1781 | active |
| thu-ml/RoboticsDiffusionTransformer RDT-1B is a 1B-parameter diffusion foundation model for robot bimanual manipulation, pre-trained on 1M+ multi-robot episodes to predict rob… | 50 | 1778 | active |
| Robbyant/lingbot-vla LingBot-VLA is a Vision-Language-Action foundation model for robot manipulation, pretrained on 20,000 hours of real-world dual-arm robot da… | 54 | 1777 | active |
| Totoro97/NeuS Official PyTorch implementation of NeuS, a neural implicit surface reconstruction method that learns SDF-based surfaces via volume renderin… | 32 | 1777 | stable |
| microsoft/MarS MarS is a financial market simulation engine powered by a generative foundation model, developed by Microsoft. It provides tools for simula… | 61 | 1776 | active |
| levihsu/OOTDiffusion Official implementation of OOTDiffusion, a latent diffusion model for controllable virtual try-on that generates images of a person wearing… | 26 | 6586 | maintenance |
| xinntao/ESRGAN ESRGAN (Enhanced SRGAN) is a PyTorch-based image super-resolution model that won the PIRM 2018 Challenge on Perceptual Super-Resolution. Th… | 32 | 6568 | maintenance |
| tsurumeso/vocal-remover A Python command-line tool that uses deep neural networks to separate vocals from instrumental tracks in songs. It includes pretrained mode… | 23 | 1757 | stable |
| SamsungSAILMontreal/TinyRecursiveModels Official codebase for the Tiny Recursive Model (TRM) paper, a recursive reasoning approach where a tiny 7M-parameter neural network iterati… | 10 | 6566 | maintenance |
| VAST-AI-Research/TripoSG TripoSG is an open-source image-to-3D generation foundation model that produces high-fidelity 3D meshes from single images using large-scal… | 27 | 1755 | active |
| microsoft/mup The `mup` Python package implements Maximal Update Parametrization (μP) for PyTorch models, enabling optimal hyperparameters to remain stab… | 23 | 1753 | stable |
| facebookresearch/metaseq Metaseq is a PyTorch codebase from Meta AI for training and working with large-scale Open Pre-trained Transformers (OPT), forked from fairs… | 10 | 6548 | maintenance |
| octo-models/octo Octo is an open-source generalist robot policy: a transformer-based diffusion policy pretrained on 800k robot trajectories from the Open X-… | 17 | 1751 | active |
| codertimo/BERT-pytorch A PyTorch implementation of Google AI's 2018 BERT model with simple, readable code. It provides CLI tools for building vocabulary and pre-t… | 23 | 6527 | maintenance |
| CompVis/taming-transformers The official implementation of 'Taming Transformers for High-Resolution Image Synthesis' (CVPR 2021), combining a convolutional VQGAN codeb… | 32 | 6521 | maintenance |
| zhouhaoyi/Informer2020 The official PyTorch implementation of Informer, an efficient Transformer architecture for long sequence time-series forecasting that won t… | 44 | 6516 | maintenance |
| TencentARC/BrushNet BrushNet is the official PyTorch implementation of an ECCV 2024 plug-and-play image inpainting model that embeds pixel-level masked image f… | 25 | 1745 | active |
| openai/consistency_models Official PyTorch implementation of Consistency Models, a generative image model family from OpenAI supporting consistency distillation, con… | 10 | 6486 | maintenance |
| google/automl Google Brain's AutoML repository containing implementations of AutoML models and libraries such as EfficientNet, EfficientNetV2, and Effici… | 10 | 6474 | maintenance |
| decisionintelligence/TFB TFB is a comprehensive and fair benchmarking framework for time series forecasting methods, covering deep learning, machine learning, and s… | 68 | 1734 | active |
| facebookresearch/multimodal TorchMultimodal is a PyTorch library from Meta for training state-of-the-art multimodal multi-task models at scale, covering both content u… | 77 | 1732 | active |
| NVIDIA/FasterTransformer NVIDIA's highly optimized C++/CUDA library for fast inference of Transformer-based models such as BERT, GPT, and encoder-decoder models, wi… | 23 | 6447 | maintenance |
| mahmoodlab/CLAM CLAM is an open-source Python toolkit for data-efficient, weakly supervised classification of whole-slide images (WSIs) in computational pa… | 39 | 1728 | active |
| facebookresearch/ConvNeXt Official PyTorch implementation of ConvNeXt, a pure convolutional neural network architecture from the CVPR 2022 paper 'A ConvNet for the 2… | 10 | 6416 | maintenance |