function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ML-GSAI/LLaDA Official PyTorch implementation of LLaDA, a family of large language diffusion models (8B base/instruct, MoE, and iLLaDA variants) with pre… | 62 | 3943 | active |
| ali-vilab/VACE VACE is the official implementation of an all-in-one video creation and editing model from Tongyi Lab, built on Wan2.1 diffusion models. It… | 41 | 3934 | active |
| thuml/Transfer-Learning-Library TLlib is a PyTorch-based open-source library for transfer learning, covering domain adaptation, task adaptation (finetuning), and domain ge… | 23 | 3931 | active |
| Tencent-Hunyuan/Hunyuan3D-2.1 Tencent's open-source 3D asset generation model that creates high-fidelity 3D meshes with production-ready PBR materials from single images… | 40 | 3917 | active |
| hustvl/Vim Vision Mamba (Vim) is a PyTorch implementation of a generic vision backbone built on bidirectional Mamba state space models, published at I… | 29 | 3899 | active |
| Plachtaa/seed-vc Seed-VC is a Python tool and model for zero-shot voice conversion, real-time voice conversion, and singing voice conversion, cloning a voic… | 10 | 3888 | active |
| safetensors/safetensors Safetensors is a simple, secure file format and library for storing and distributing tensors, designed as a fast zero-copy alternative to p… | 91 | 3876 | stable |
| NVlabs/VILA VILA is a family of open-source vision language models (VLMs) optimized for efficient video and multi-image understanding, spanning edge, d… | 57 | 3857 | active |
| open-mmlab/mmpretrain MMPretrain is OpenMMLab's PyTorch-based toolbox and benchmark for image classification model pre-training, covering supervised, self-superv… | 23 | 3850 | active |
| Stability-AI/stable-audio-tools Stability AI's training and inference toolkit for conditional audio generation models, including Stable Audio Open. It supports training cu… | 73 | 3849 | active |
| shenweichen/GraphEmbedding A Python library providing implementations of classic graph embedding algorithms including DeepWalk, LINE, Node2Vec, SDNE, and Struc2Vec. I… | 68 | 3845 | active |
| neuraloperator/neuraloperator A PyTorch library for learning neural operators, which map between function spaces rather than finite-dimensional vectors. It provides the … | 82 | 3836 | active |
| google-research/scenic Scenic is a JAX-based library from Google Research focused on attention-based models for computer vision, providing shared lightweight libr… | 76 | 3821 | active |
| ddbourgin/numpy-ml numpy-ml is a collection of machine learning models and algorithms implemented exclusively in NumPy and the Python standard library, coveri… | 32 | 16330 | maintenance |
| lightly-ai/lightly LightlySSL is a Python library built on PyTorch for self-supervised learning on images, offering modular implementations of methods like Si… | 93 | 3797 | active |
| google/deepvariant DeepVariant is a deep learning-based genomic variant caller that converts aligned DNA sequencing reads (BAM/CRAM) into pileup image tensors… | 70 | 3791 | stable |
| SandAI-org/MAGI-1 MAGI-1 is an open-source autoregressive video generation model from Sand.ai, released with Apache-2.0 licensed code and weights. It generat… | 59 | 3772 | active |
| HeartMuLa/heartlib HeartMuLa is a family of open-source music foundation models that generate music conditioned on lyrics and tags with multilingual support. … | 50 | 3749 | active |
| fudan-generative-vision/hallo2 Hallo2 is a Python research library from Fudan University that animates a single portrait image using audio input, producing long-duration … | 26 | 3734 | active |
| genmoai/mochi Mochi 1 is Genmo's open-source, state-of-the-art text-to-video generation model released under Apache 2.0, with a Python API, CLI, and Grad… | 46 | 3713 | active |
| microsoft/Bringing-Old-Photos-Back-to-Life The official PyTorch implementation of 'Bringing Old Photos Back to Life' (CVPR 2020 Oral), a deep learning model that restores old photos … | 23 | 15704 | maintenance |
| facebookresearch/map-anything MapAnything is an open-source research framework from Meta and CMU for universal feed-forward metric 3D reconstruction using an end-to-end … | 77 | 3682 | active |
| ferdous-alam/GenCAD GenCAD is a research codebase for image-conditioned CAD model generation using transformer-based contrastive representations (CCIP) and dif… | 36 | 3669 | active |
| HazyResearch/ThunderKittens ThunderKittens is a C++/CUDA framework of tile-based primitives for writing fast deep learning GPU kernels. It embeds natively into CUDA so… | 70 | 3659 | active |
| shitagaki-lab/see-through A research framework from a SIGGRAPH 2026 paper that decomposes a single anime character illustration into up to 23 fully inpainted, semant… | 58 | 3641 | active |
| kaldi-asr/kaldi Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea… | 52 | 15469 | maintenance |
| opendilab/DI-engine DI-engine is an open-source reinforcement learning framework from OpenDILab that provides comprehensive implementations of deep RL algorith… | 48 | 3638 | active |
| cmusatyalab/openface OpenFace is a free and open source Python and Torch implementation of face recognition based on Google's FaceNet deep neural network. It ge… | 65 | 15438 | maintenance |
| facebookresearch/detr DETR is Facebook Research's PyTorch implementation of Detection Transformer, an end-to-end object detection model that replaces hand-crafte… | 10 | 15354 | maintenance |
| facebookresearch/sam-audio SAM-Audio is Meta's foundation model for isolating any sound in audio using text, visual, or temporal prompts. This repository provides inf… | 55 | 3612 | active |
| albumentations-team/albumentations Albumentations is a fast, flexible Python image augmentation library for computer vision, supporting images, masks, bounding boxes, keypoin… | 10 | 15315 | maintenance |
| AI4Finance-Foundation/FinRL-Trading FinRL-X is an open-source, AI-native modular infrastructure for quantitative trading that unifies data processing, strategy composition, ba… | 71 | 3592 | active |
| ZhaoJ9014/face.evoLVe A high-performance face recognition library built on PaddlePaddle and PyTorch, providing comprehensive tools for face-related analytics and… | 38 | 3589 | active |
| AiuniAI/Unique3D Unique3D is the official implementation of a NeurIPS 2024 paper that generates high-quality textured 3D meshes from a single image in about… | 38 | 3579 | active |
| GVCLab/PersonaLive PersonaLive is a diffusion-based framework for real-time, streamable portrait image animation, generating infinite-length expressive talkin… | 53 | 3552 | active |
| sdv-dev/SDV SDV (Synthetic Data Vault) is a Python library for generating synthetic tabular data using machine learning models ranging from GaussianCop… | 99 | 3549 | stable |
| AliaksandrSiarohin/first-order-model Official PyTorch/Jupyter implementation of the First Order Motion Model for image animation (NeurIPS 2019). It animates a static source ima… | 32 | 15015 | maintenance |
| starVLA/starVLA StarVLA is an open-source, Lego-like modular codebase for developing Vision-Language-Action (VLA) models for generalist robots. It unifies … | 69 | 3530 | active |
| google-research/big_vision Google Research's official Jax/Flax codebase for training large-scale vision models such as Vision Transformer, SigLIP, MLP-Mixer, and LiT … | 42 | 3528 | active |
| cszn/KAIR A PyTorch image restoration toolbox providing training and testing code for many restoration models including DnCNN, FFDNet, SRMD, USRNet, … | 23 | 3523 | active |
| neonbjb/tortoise-tts Tortoise TTS is a multi-voice text-to-speech library built on PyTorch that prioritizes highly realistic prosody and intonation. It combines… | 32 | 14870 | maintenance |
| pathwaycom/bdh BDH (Dragon Hatchling) is a biologically inspired large language model architecture that bridges deep learning and neuroscience, implemente… | 54 | 3519 | active |
| NVlabs/FoundationPose FoundationPose is NVIDIA's unified foundation model for 6D object pose estimation and tracking of novel objects, supporting both model-base… | 62 | 3516 | active |
| MooreThreads/Moore-AnimateAnyone An open-source reproduction of AnimateAnyone that animates a character from a single reference image using pose sequences from a driving vi… | 26 | 3514 | active |
| google-deepmind/alphafold Open-source implementation of the AlphaFold 2 inference pipeline for predicting protein structures from amino acid sequences, including Alp… | 58 | 14811 | maintenance |
| NVIDIA/TransformerEngine Transformer Engine is an NVIDIA library for accelerating Transformer model training and inference on NVIDIA GPUs using low-precision format… | 99 | 3504 | active |
| PKU-YuanGroup/Video-LLaVA Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into… | 27 | 3500 | active |
| autorope/donkeycar Donkeycar is an open-source Python library and hardware platform for building small-scale self-driving RC cars with Raspberry Pi or Jetson … | 88 | 3495 | active |
| borisdayma/dalle-mini DALL·E Mini is a Python library and model that generates images from a text prompt, available via pip and hosted on Hugging Face Model Hub.… | 23 | 14740 | maintenance |
| facebookresearch/ijepa Official PyTorch implementation of I-JEPA, a self-supervised learning method that predicts latent representations of image regions from oth… | 10 | 3489 | active |
| NVIDIA/Model-Optimizer NVIDIA Model Optimizer (ModelOpt) is a Python library of state-of-the-art model optimization techniques including quantization, pruning, di… | 91 | 3488 | active |
| guandeh17/Self-Forcing Official implementation of Self Forcing, a training method for autoregressive video diffusion models that simulates inference during traini… | 37 | 3488 | active |
| POSTECH-CVLab/PyTorch-StudioGAN PyTorch-StudioGAN is a PyTorch library providing unified implementations of representative GAN architectures (BigGAN, StyleGAN2/3, etc.) fo… | 23 | 3487 | stable |
| Tencent-Hunyuan/Hunyuan3D-1 Tencent Hunyuan3D-1.0 is an open-source two-stage diffusion-based model for generating 3D assets from text prompts or images. It provides i… | 46 | 3482 | active |
| MiniMax-AI/MiniMax-01 Official repository for MiniMax-Text-01 and MiniMax-VL-01, open-weight large language and vision-language models built on a linear attentio… | 34 | 3466 | active |
| NVlabs/Eagle Eagle is NVIDIA's family of frontier vision-language models (Eagle, Eagle 2, Eagle 2.5) built with data-centric training strategies, plus L… | 64 | 3462 | active |
| shenweichen/DeepCTR-Torch DeepCTR-Torch is a PyTorch library providing easy-to-use, modular, and extendable implementations of deep-learning-based CTR (click-through… | 78 | 3450 | active |
| XinJingHao/DRL-Pytorch A unified PyTorch implementation collection of popular deep reinforcement learning algorithms including DQN variants, PPO, DDPG, TD3, SAC, … | 44 | 3436 | active |
| NVlabs/stylegan The official TensorFlow implementation of StyleGAN, NVIDIA's style-based generator architecture for generative adversarial networks from th… | 32 | 14416 | maintenance |
| aqlaboratory/openfold OpenFold is a faithful, trainable PyTorch reproduction of DeepMind's AlphaFold 2 for protein structure prediction. It is memory-efficient a… | 48 | 3420 | active |
| microsoft/nni NNI (Neural Network Intelligence) is an open-source AutoML toolkit from Microsoft that automates hyperparameter tuning, neural architecture… | 10 | 14361 | maintenance |
| davidsandberg/facenet A TensorFlow implementation of the FaceNet face recognizer that generates 128-dimensional face embeddings, including face detection via MTC… | 32 | 14343 | maintenance |
| WongKinYiu/yolov7 Official PyTorch implementation of the YOLOv7 paper, a state-of-the-art real-time object detector with trainable bag-of-freebies techniques… | 23 | 14139 | maintenance |
| CompVis/latent-diffusion The official research code and pretrained model zoo for Latent Diffusion Models (LDM), the paper behind Stable Diffusion, enabling high-res… | 32 | 14133 | maintenance |
| nv-tlabs/kimodo Kimodo is NVIDIA's official implementation of a kinematic motion diffusion model trained on 700 hours of motion capture data to generate hi… | 56 | 3365 | active |
| OpenTalker/SadTalker SadTalker is a CVPR 2023 deep learning tool that generates realistic talking head videos from a single portrait image and an audio clip by … | 22 | 14040 | maintenance |
| VainF/Torch-Pruning Torch-Pruning is a PyTorch framework for structural neural network pruning based on the DepGraph algorithm from CVPR 2023. It automatically… | 50 | 3348 | active |
| magenta/ddsp DDSP is a Python library of differentiable digital signal processing components (synthesizers, filters, waveshapers) that can be embedded i… | 64 | 3344 | active |
| opengeos/geoai GeoAI is a Python package that integrates artificial intelligence with geospatial data analysis, built on PyTorch, Transformers, and segmen… | 89 | 3327 | active |
| Peterande/D-FINE D-FINE is the official PyTorch implementation of an ICLR 2025 Spotlight paper that redefines the regression task in DETR-style detectors as… | 67 | 3305 | active |
| microsoft/LoRA loralib is the official PyTorch implementation of LoRA (Low-Rank Adaptation), which fine-tunes large language models by injecting trainable… | 23 | 13767 | maintenance |
| Tencent-Hunyuan/HunyuanImage-3.0 HunyuanImage-3.0 is Tencent's open-source native multimodal model for text-to-image and image-to-image generation, with inference code and … | 57 | 3253 | active |
| Beckschen/TransUNet Official PyTorch implementation of TransUNet, a U-Net-style architecture that uses a Vision Transformer encoder for medical image segmentat… | 63 | 3234 | stable |
| mit-han-lab/bevfusion BEVFusion is a PyTorch-based multi-task multi-sensor fusion framework that unifies camera and LiDAR features in a shared bird's-eye view re… | 10 | 3230 | stable |
| Jittor/jittor Jittor is a high-performance deep learning framework from Tsinghua University based on just-in-time (JIT) compilation and meta-operators, w… | 67 | 3229 | active |
| onnx/onnx-tensorrt A C++ parser library and backend that converts ONNX models into TensorRT engines for high-performance GPU inference. It is maintained by NV… | 92 | 3228 | active |
| MzeroMiko/VMamba VMamba is a PyTorch implementation of a visual state space model (SSM) vision backbone based on Mamba, featuring 2D Selective Scan (SS2D) f… | 21 | 3219 | active |
| kerlomz/captcha_trainer A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren… | 55 | 3213 | active |
| jy0205/Pyramid-Flow Pyramid Flow is the official PyTorch implementation of a training-efficient autoregressive video generation model based on pyramidal flow m… | 22 | 3208 | active |
| NVIDIA/physicsnemo NVIDIA PhysicsNeMo is an open-source Python deep-learning framework for building, training, fine-tuning, and inferring physics AI models us… | 89 | 3198 | active |
| Pointcept Pointcept is a PyTorch-based research codebase for point cloud perception, providing implementations of state-of-the-art 3D scene understan… | 76 | 3196 | active |
| facebookresearch/dinov2 PyTorch implementation and pretrained models for DINOv2, a self-supervised vision transformer method from Meta AI that learns robust visual… | 68 | 13266 | maintenance |
| LeelaChessZero/lc0 Lc0 is an open-source, UCI-compliant chess engine that plays chess using neural networks trained via AlphaZero-style self-play reinforcemen… | 65 | 3193 | active |
| stepfun-ai/Step-Video-T2V Step-Video-T2V is an open-source text-to-video generation model from StepFun, released with inference code and pretrained weights (includin… | 25 | 3187 | active |
| MiniMax-AI/MiniMax-M1 MiniMax-M1 is an open-weight, large-scale hybrid-attention reasoning language model released by MiniMax under Apache-2.0. The repository pr… | 32 | 3180 | active |
| Rudrabha/Wav2Lip Wav2Lip is the official research code for the ACM Multimedia 2020 paper 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In … | 45 | 13182 | maintenance |
| facebookresearch/tribev2 TRIBE v2 is a multimodal deep learning model from Meta AI that predicts fMRI brain responses to naturalistic video, audio, and text stimuli… | 55 | 3172 | active |
| ali-vilab/VGen VGen is the official repository for a holistic video generation ecosystem built on diffusion models, including the I2VGen-XL cascaded image… | 27 | 3155 | active |
| megvii-research/NAFNet NAFNet is the official PyTorch implementation of a state-of-the-art image restoration network that removes nonlinear activation functions. … | 32 | 3148 | stable |
| junyanz/CycleGAN A Torch (Lua) implementation of CycleGAN and pix2pix for unpaired image-to-image translation using cycle-consistent adversarial networks. I… | 32 | 12870 | maintenance |
| thuml/Time-Series-Library TSLib is an open-source Python library providing a unified codebase of advanced deep learning models for general time series analysis. It s… | 66 | 12785 | maintenance |
| guillaume-be/rust-bert A Rust-native library providing ready-to-use NLP pipelines and transformer-based models (BERT, DistilBERT, GPT-2, RoBERTa, BART, etc.), por… | 60 | 3076 | active |
| imbue-bit/AlphaGPT AlphaGPT is an open-source automated factor factory based on deep reinforcement learning for quantitative finance. It mines and generates a… | 55 | 3073 | active |
| tensorflow/tflite-micro TensorFlow Lite for Microcontrollers (TFLM) is a C++ port of TensorFlow Lite for running ML models on microcontrollers, DSPs, and other mem… | 77 | 3059 | active |
| RosettaCommons/RFdiffusion RFdiffusion is an open-source method for de novo protein structure generation using diffusion models, with or without conditional informati… | 62 | 3026 | active |
| deepseek-ai/DreamCraft3D Official PyTorch implementation of DreamCraft3D, an ICLR 2024 hierarchical 3D content generation method that turns a single 2D image into a… | 35 | 3021 | stable |
| osmr/imgclsmob A research sandbox providing (re)implementations of numerous deep learning computer vision models for classification, segmentation, detecti… | 23 | 3016 | active |
| ZQPei/deep_sort_pytorch A PyTorch implementation of the Deep SORT multi-object tracking algorithm, pairing YOLOv3/YOLOv5 (or Mask R-CNN) detectors with a CNN re-id… | 32 | 3012 | active |
| benedekrozemberczki/pytorch_geometric_temporal PyTorch Geometric Temporal is a temporal (dynamic) extension library for PyTorch Geometric providing spatiotemporal signal processing with … | 68 | 2992 | active |
| MeiGen-AI/MultiTalk MultiTalk is an audio-driven framework for generating multi-person conversational videos from multi-stream audio, a reference image, and a … | 56 | 2992 | active |