function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| TorchIO-project/torchio TorchIO is a Python library for loading, augmenting, and processing 3D medical images (MRI, CT) within PyTorch deep learning pipelines. It … | 93 | 2439 | active |
| chemprop/chemprop Chemprop is a Python package implementing message passing neural networks (MPNNs) for predicting molecular and reaction properties. It prov… | 94 | 2438 | active |
| shenweichen/DeepMatch DeepMatch is a Python library of deep matching models for recommendations and advertising, built on TensorFlow/Keras. It lets users train m… | 72 | 2433 | active |
| tflearn/tflearn TFLearn is a modular deep learning library providing a higher-level, Keras-like API on top of TensorFlow for building and training neural n… | 23 | 9576 | maintenance |
| wolny/pytorch-3dunet A PyTorch implementation of 3D U-Net and its variants (residual, squeeze-and-excitation) for volumetric semantic segmentation, with 2D U-Ne… | 63 | 2416 | active |
| HuCaoFighting/Swin-Unet Official PyTorch implementation of Swin-Unet, a U-shaped pure Transformer model for medical image segmentation, published at ECCV 2022 Medi… | 41 | 2416 | stable |
| X-PLUG/mPLUG-DocOwl mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO… | 39 | 2411 | active |
| AI-Hypercomputer/maxtext MaxText is a high-performance, scalable open-source LLM training library written in pure Python/JAX, targeting Google Cloud TPUs and GPUs. … | 97 | 2408 | active |
| resemble-ai/resemble-enhance Resemble Enhance is an AI-powered Python tool that improves speech quality through denoising and enhancement, using a denoiser module and a… | 17 | 2397 | active |
| Alibaba-Quark/LiveAvatar LiveAvatar is an open-source implementation of an ECCV 2026 paper for streaming, real-time, infinite-length audio-driven avatar video gener… | 61 | 2386 | active |
| google/neural-tangents Neural Tangents is a Python library built on JAX for defining, training, and evaluating neural networks of both finite and infinite width. … | 10 | 2383 | stable |
| apple/axlearn AXLearn is a Python deep learning library built on JAX and XLA for developing and training large-scale models, with an object-oriented conf… | 71 | 2372 | active |
| zai-org/GLM-V GLM-V is the open-source repository for Zhipu AI's GLM-4.6V, GLM-4.5V, and GLM-4.1V-Thinking vision-language models, which perform versatil… | 60 | 2370 | active |
| OpenGVLab/InternVideo InternVideo is a series of open-source video foundation models for multimodal video understanding, spanning generative and discriminative l… | 72 | 2368 | active |
| tencent-ailab/V-Express V-Express is a Python research project from Tencent AI Lab that generates talking head portrait videos from a reference image, audio, and V… | 25 | 2360 | active |
| alibaba/EasyRec EasyRec is a TensorFlow-based framework from Alibaba for building large-scale deep learning recommendation models covering matching, rankin… | 65 | 2356 | active |
| facebookresearch/perception_models Meta's Perception Models repository hosting state-of-the-art image, video, and audio encoders (Perception Encoder, PE) and a multimodal lan… | 54 | 2353 | active |
| mlech26l/ncps A Python package providing PyTorch and TensorFlow/Keras implementations of Neural Circuit Policies (NCPs), including liquid time-constant (… | 23 | 2345 | active |
| ASLP-lab/DiffRhythm DiffRhythm is an open-source latent diffusion model for end-to-end full-length song generation, supporting lyrics-to-song, text style promp… | 44 | 2339 | active |
| google-deepmind/optax Optax is a gradient processing and optimization library for JAX, offering composable building blocks like optimizers and loss functions. It… | 82 | 2325 | stable |
| PKU-YuanGroup/MoE-LLaVA MoE-LLaVA is an open-source Mixture-of-Experts based sparse large vision-language model, released with the MoE-Tuning training strategy fro… | 31 | 2322 | active |
| Cadene/pretrained-models.pytorch A Python library providing pretrained ConvNet models (ResNet, ResNeXt, InceptionV4, Xception, NASNet, SENet, DPN, etc.) for PyTorch behind … | 32 | 9099 | maintenance |
| SkyworkAI/Matrix-Game Matrix-Game is Skywork AI's open-source series of interactive world foundation models that generate real-time, streaming video in response … | 52 | 2314 | active |
| facebookresearch/ImageBind A PyTorch library from Meta AI implementing ImageBind, a model that learns a joint embedding space across six modalities: images, text, aud… | 54 | 9064 | maintenance |
| frgfm/torch-cam TorchCAM is a Python library that extracts class activation maps (CAMs) from PyTorch CNN classifiers, supporting many CAM variants such as … | 72 | 2304 | active |
| YvanYin/Metric3D Metric3D is the official PyTorch implementation of Metric3Dv1 and Metric3Dv2, monocular geometric foundation models that predict metric dep… | 33 | 2302 | active |
| NVlabs/nvdiffrec nvdiffrec is NVIDIA's official implementation of a CVPR 2022 oral paper that jointly optimizes triangular 3D meshes, PBR materials, and lig… | 76 | 2296 | stable |
| OpenHelix-Team/VLA-Adapter VLA-Adapter is the official implementation of a tiny-scale vision-language-action model that bridges vision-language representations to rob… | 50 | 2294 | active |
| traveller59/spconv SpConv is a spatially sparse convolution library for deep learning on 3D point clouds and sparse tensors, distributed as PyPI packages with… | 32 | 2291 | active |
| huggingface/picotron Picotron is a minimalist, hackable distributed training framework for pre-training Llama-like large language models using 4D parallelism (d… | 40 | 2289 | active |
| codeplea/genann Genann is a minimal, well-tested C99 library for training and running feedforward artificial neural networks. It is contained in a single s… | 97 | 2280 | stable |
| aigc-apps/EasyAnimate EasyAnimate is an end-to-end Python pipeline for high-resolution, long video and image generation based on transformer diffusion (DiT) mode… | 20 | 2270 | active |
| 666DZY666/micronet micronet is a Python library for deep neural network model compression and deployment built on PyTorch. It provides quantization (QAT, PTQ,… | 41 | 2266 | active |
| ashawkey/stable-dreamfusion A PyTorch implementation of Dreamfusion that generates 3D models from text prompts or images using NeRF combined with Stable Diffusion guid… | 23 | 8854 | maintenance |
| hkchengrex/MMAudio MMAudio is a PyTorch-based model for generating synchronized audio from video and/or text inputs, using multimodal joint training across au… | 43 | 2264 | active |
| nv-tlabs/lyra Project Lyra is NVIDIA's open series of generative 3D world models, including Lyra 1.0 for feed-forward 3D/4D scene generation from a singl… | 59 | 2262 | active |
| opendatalab/DocLayout-YOLO DocLayout-YOLO is a real-time YOLO-v10-based model for detecting document layout elements (text blocks, tables, figures, etc.) in diverse d… | 29 | 2258 | active |
| RosettaCommons/RoseTTAFold RoseTTAFold is the official implementation of a deep learning system for predicting protein structures and interactions using a three-track… | 23 | 2258 | stable |
| stepfun-ai/Step1X-Edit Step1X-Edit is an open-source state-of-the-art instruction-based image editing model from StepFun, designed to rival closed-source editors … | 55 | 2256 | active |
| fishaudio/Bert-VITS2 Bert-VITS2 is a text-to-speech model implementation combining the VITS2 architecture with multilingual BERT embeddings, written in Python. … | 63 | 8796 | maintenance |
| CoinCheung/pytorch-loss A PyTorch library providing a collection of loss functions (focal loss, triplet loss, AMSoftmax, label-smooth CE, dice loss, lovasz-softmax… | 32 | 2252 | active |
| facebookresearch/fvcore fvcore is a lightweight Python core library providing common functionality shared across FAIR's computer vision frameworks such as Detectro… | 76 | 2250 | stable |
| Alpha-VLLM/Lumina-T2X Lumina-T2X is a unified framework for text-to-any-modality generation built on flow-based large diffusion transformers. It supports generat… | 28 | 2250 | active |
| azavea/raster-vision Raster Vision is an open source Python library and low-code framework for building computer vision models on satellite, aerial, and other l… | 61 | 2240 | active |
| facebookresearch/fairchem FAIR Chemistry's centralized Python library of machine learning models, datasets, and applications for materials science and quantum chemis… | 99 | 2231 | active |
| microsoft/LLaVA-Med LLaVA-Med is a large language-and-vision assistant fine-tuned for the biomedicine domain, built on the LLaVA multimodal architecture. It su… | 40 | 2231 | active |
| NVIDIA/vid2vid A PyTorch implementation of NVIDIA's video-to-video synthesis method for generating high-resolution (e.g., 2048x1024) photorealistic videos… | 32 | 8692 | maintenance |
| facebookresearch/DiT Official PyTorch implementation of Diffusion Transformers (DiT) from the paper 'Scalable Diffusion Models with Transformers', including mod… | 10 | 8689 | maintenance |
| apple/ml-ferret Apple's Ferret, an end-to-end multimodal large language model (MLLM) that accepts any-form referring and grounds anything in its responses,… | 27 | 8674 | maintenance |
| NVlabs/MambaVision MambaVision is NVIDIA's official PyTorch implementation of a hybrid Mamba-Transformer vision backbone, published at CVPR 2025. It provides … | 49 | 2224 | active |
| ellisdg/3DUnetCNN A PyTorch library for building, training, and applying 3D U-Net convolutional neural networks for medical image segmentation. It provides c… | 45 | 2224 | active |
| ahmedfgad/GeneticAlgorithmPython PyGAD is an open-source Python 3 library for implementing the genetic algorithm to optimize single- and multi-objective problems. It can al… | 80 | 2220 | active |
| lifeiteng/vall-e An unofficial PyTorch implementation of VALL-E, a zero-shot text-to-speech model that treats TTS as a conditional language modeling task ov… | 40 | 2215 | active |
| aigc-apps/VideoX-Fun VideoX-Fun is a Python-based video generation pipeline built on Diffusion Transformer models (CogVideoX-Fun, Wan-Fun) that generates videos… | 67 | 2210 | active |
| RubixML/ML Rubix ML is a high-level open-source machine learning and deep learning library for PHP, offering 40+ supervised and unsupervised algorithm… | 99 | 2204 | active |
| thuml/iTransformer Official PyTorch implementation of iTransformer, an ICLR 2024 Spotlight paper that inverts the Transformer architecture for multivariate ti… | 41 | 2200 | stable |
| lucidrains/lion-pytorch A PyTorch implementation of the Lion optimizer (Evolved Sign Momentum), discovered by Google Brain via genetic algorithms and claimed to ou… | 62 | 2199 | active |
| NX-AI/xlstm Official PyTorch implementation of xLSTM, an extended Long Short-Term Memory recurrent architecture with exponential gating and matrix memo… | 59 | 2198 | active |
| Harry24k/adversarial-attacks-pytorch Torchattacks is a PyTorch library providing implementations of adversarial attacks to generate adversarial examples against deep learning m… | 23 | 2177 | active |
| ByteDance-Seed/VeOmni VeOmni is a PyTorch-native framework for single- and multi-modal model pre-training and post-training, with a modular, trainer-free design … | 82 | 2173 | active |
| facebookresearch/mae A PyTorch/GPU re-implementation of the Masked Autoencoders (MAE) paper for self-supervised vision learning. It includes pre-training code, … | 10 | 8370 | maintenance |
| galilai-group/stable-worldmodel A Python library providing a unified platform for reproducible world model research, covering data collection, training, and evaluation via… | 81 | 2156 | active |
| TencentARC/Pixal3D Pixal3D is a research codebase for generating high-fidelity 3D assets from a single image using a pixel-aligned generation paradigm that ba… | 54 | 2156 | active |
| jd-opensource/JoyAI-Image JoyAI-Image is a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editin… | 58 | 2148 | active |
| google/trax Trax is an end-to-end deep learning library built on JAX and TensorFlow that focuses on clear code and speed, developed and maintained by t… | 10 | 8306 | maintenance |
| ViTAE-Transformer/ViTPose Official PyTorch implementation of ViTPose and ViTPose++, Vision Transformer models for human and generic body pose estimation from NeurIPS… | 59 | 2138 | stable |
| PKU-YuanGroup/LLaVA-CoT LLaVA-CoT is an 11B vision language model with training, inference, and dataset-generation code for spontaneous step-by-step multimodal rea… | 47 | 2132 | active |
| utkuozbulak/pytorch-cnn-visualizations A PyTorch library implementing a wide range of convolutional neural network visualization and interpretability techniques, including Grad-C… | 32 | 8233 | maintenance |
| lukemelas/EfficientNet-PyTorch A PyTorch implementation of the EfficientNet convolutional neural network family with pretrained ImageNet weights. It provides a simple pip… | 23 | 8222 | maintenance |
| yyfz/Pi3 Pi3 (π³) is a feed-forward neural network for visual geometry reconstruction that eliminates the need for a fixed reference view, using a p… | 59 | 2122 | active |
| autonomousvision/sdfstudio SDFStudio is a unified and modular framework for neural implicit surface reconstruction built on top of nerfstudio. It provides unified imp… | 31 | 2120 | active |
| GAIR-NLP/daVinci-MagiHuman daVinci-MagiHuman is an open-source 15B-parameter single-stream transformer foundation model that jointly generates synchronized audio and … | 49 | 2113 | active |
| 3DTopia/LGM LGM is the official PyTorch implementation of an ECCV 2024 Oral paper that generates high-resolution 3D models from text prompts or single-… | 26 | 2111 | active |
| fangwei123456/spikingjelly SpikingJelly is an open-source deep learning framework for Spiking Neural Networks (SNNs) built on PyTorch. It provides a beginner-friendly… | 77 | 2110 | active |
| River-Zhang/ICEdit ICEdit (In-Context Edit) is a research framework for instruction-based image editing built on large-scale Diffusion Transformers, using a L… | 45 | 2102 | active |
| eloialonso/diamond DIAMOND is a Python implementation of a reinforcement learning agent trained entirely inside a diffusion-based world model, presented as a … | 24 | 2096 | active |
| SUDO-AI-3D/zero123plus Zero123++ is a diffusion base model that generates consistent multi-view images from a single input image, intended as a stepping stone for… | 27 | 2095 | active |
| ContinualAI/avalanche Avalanche is an end-to-end continual learning library built on PyTorch, developed by ContinualAI. It provides modules for benchmarks, train… | 27 | 2088 | active |
| alex-damian/pulse PULSE is a Python research implementation of a CVPR 2020 paper that upscales low-resolution face photos by searching the latent space of a … | 32 | 8023 | maintenance |
| facebookresearch/ai4animationpy AI4AnimationPy is a Python framework for AI-driven character animation using neural networks, providing motion capture processing, training… | 59 | 2078 | active |
| PKU-YuanGroup/Helios Helios is a 14B autoregressive diffusion model for real-time, minute-scale video generation supporting text-to-video, image-to-video, and v… | 59 | 2076 | active |
| facebookresearch/ConvNeXt-V2 Official PyTorch implementation of ConvNeXt V2, a family of pure convolutional neural network models co-designed with a fully convolutional… | 10 | 2069 | stable |
| brightmart/text_classification A collection of deep learning baseline models for text classification in NLP, implemented in TensorFlow. It covers classic architectures li… | 32 | 7940 | maintenance |
| Plachtaa/VALL-E-X An open-source Python implementation of Microsoft's VALL-E X zero-shot text-to-speech model, with a community-trained pretrained checkpoint… | 10 | 7931 | maintenance |
| shanglianlm0525/PyTorch-Networks A collection of PyTorch implementations of classic and modern CNN architectures, covering classification, detection, segmentation, face, an… | 53 | 2055 | active |
| facebookresearch/theseus Theseus is a PyTorch-based library for building custom differentiable nonlinear optimization layers, supporting problems in robotics and vi… | 23 | 2055 | active |
| jaywalnut310/vits VITS is the official PyTorch implementation of an end-to-end text-to-speech model based on a conditional variational autoencoder with adver… | 32 | 7889 | maintenance |
| WenjieDu/PyPOTS PyPOTS is a Python toolbox for machine learning and data mining on partially-observed time series with missing values. It integrates 50+ st… | 93 | 2051 | active |
| bytedance/Protenix Protenix is an open-source PyTorch reproduction of AlphaFold 3 for high-accuracy biomolecular structure prediction, covering proteins, nucl… | 79 | 2038 | active |
| jeshraghian/snntorch snnTorch is a Python library for gradient-based deep learning with spiking neural networks, built as an extension of PyTorch. It provides s… | 67 | 2036 | active |
| deep-floyd/IF DeepFloyd IF is an open-source text-to-image model library implementing a cascaded pixel diffusion architecture with a frozen T5 text encod… | 22 | 7804 | maintenance |
| pykeen/pykeen PyKEEN is a Python library for training and evaluating multimodal knowledge graph embedding models built on PyTorch. It provides a high-lev… | 70 | 2032 | active |
| tensorflow/recommenders TensorFlow Recommenders is a Python library for building recommender system models on top of TensorFlow and Keras. It covers the full workf… | 84 | 2027 | active |
| NUS-HPC-AI-Lab/VideoSys VideoSys is an open-source Python library providing easy and efficient infrastructure for video generation, supporting training, inference,… | 45 | 2022 | active |
| deepmodeling/deepmd-kit DeePMD-kit is a deep learning package for building many-body potential energy representations and running molecular dynamics simulations. I… | 95 | 2021 | stable |
| SakanaAI/continuous-thought-machines The Continuous Thought Machine (CTM) is a neural network architecture from Sakana AI that uses neuron-level temporal dynamics and neural sy… | 46 | 2019 | active |
| NVlabs/SPADE Official PyTorch implementation of SPADE (GauGAN), a CVPR 2019 method for synthesizing photorealistic images from semantic segmentation map… | 32 | 7717 | maintenance |
| tdrussell/diffusion-pipe A Python training script for fine-tuning diffusion models (image and video generation) using DeepSpeed pipeline parallelism across multiple… | 67 | 2015 | active |
| cambrian-mllm/cambrian Cambrian-1 is a fully open family of vision-centric multimodal large language models (MLLMs) from NYU's VISIONx group, with training and ev… | 47 | 2013 | active |
| philipperemy/keras-tcn A Keras/TensorFlow implementation of Temporal Convolutional Networks (TCN) with dilated causal convolutions, usable as a drop-in layer alte… | 62 | 2012 | active |