domain: deep-learning
2771 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| pq-yang/MatAnyone MatAnyone is a CVPR 2025 human video matting framework that extracts alpha mattes of target people from video using consistent memory propa… | 54 | 1605 | active |
| meituan/YOLOv6 YOLOv6 is a single-stage object detection framework implemented in PyTorch, designed for industrial applications with a family of pretraine… | 23 | 5895 | maintenance |
| ml4a/ml4a ml4a is a Python library and collection of Jupyter notebooks for making art with machine learning. It wraps popular deep learning models li… | 32 | 1602 | active |
| autonomousvision/transfuser Official PyTorch implementation of TransFuser, a transformer-based multi-modal sensor fusion model for end-to-end autonomous driving, publi… | 53 | 1599 | stable |
| X-LANCE/AniTalker AniTalker is the official PyTorch implementation of an ACM MM 2024 paper that animates a single static portrait into a vivid talking-face v… | 24 | 1598 | active |
| SonyResearch/micro_diffusion Official implementation of Sony Research's micro-budget approach to training large-scale text-to-image diffusion transformer models from sc… | 24 | 1594 | active |
| facebookresearch/fast3r Fast3R is the official PyTorch implementation of a CVPR 2025 model from Meta FAIR that reconstructs 3D scenes and estimates camera poses fr… | 10 | 1593 | active |
| google/uncertainty-baselines A library of high-quality, minimal-dependency implementations of standard and state-of-the-art uncertainty and robustness methods for deep … | 77 | 1592 | active |
| alibaba/Pai-Megatron-Patch Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and V… | 56 | 1591 | active |
| Tencent-Hunyuan/HunyuanWorld-Voyager HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ… | 52 | 1590 | active |
| intel/auto-round AutoRound is Intel's advanced quantization toolkit for LLMs and vision-language models, using sign-gradient descent to achieve high accurac… | 91 | 1588 | active |
| databricks/megablocks MegaBlocks is a lightweight Python library for efficient training of mixture-of-experts (MoE) models, built around its dropless-MoE (dMoE) … | 62 | 1588 | active |
| MoonshotAI/Kimi-Linear Kimi Linear is a hybrid linear attention architecture (Kimi Delta Attention, based on Gated DeltaNet) released by Moonshot AI with 48B-para… | 40 | 1588 | active |
| FoundationVision/Infinity Infinity is a bitwise autoregressive text-to-image generation model (CVPR 2025 Oral) with released training and inference code, checkpoints… | 56 | 1587 | active |
| Drexubery/ViewCrafter ViewCrafter is a research codebase that uses video diffusion models to synthesize high-fidelity novel views of scenes from a single or spar… | 49 | 1587 | active |
| DeepGraphLearning/torchdrug TorchDrug is a PyTorch-based machine learning library for drug discovery, covering graph neural networks, deep generative models, and reinf… | 23 | 1587 | active |
| yakhyo/uniface UniFace is a unified Python library for face analysis that bundles detection, recognition, landmark localization, face parsing, gaze estima… | 88 | 1586 | active |
| Tencent-Hunyuan/HY-WorldPlay HY-WorldPlay (HY-World 1.5) is Tencent Hunyuan's open-source framework for interactive 3D world modeling, generating explorable 3D scenes f… | 55 | 1586 | active |
| pytorch/FBGEMM FBGEMM is a collection of highly optimized low-precision matrix multiplication and convolution kernels for server-side deep learning infere… | 93 | 1584 | active |
| XPixelGroup/HAT HAT (Hybrid Attention Transformer) is a PyTorch implementation of a state-of-the-art transformer model for image super-resolution and resto… | 32 | 1583 | stable |
| microsoft/MMdnn MMdnn is a Microsoft toolkit for converting, visualizing, and diagnosing deep learning models across frameworks such as TensorFlow, PyTorch… | 39 | 5805 | maintenance |
| ai-forever/ghost GHOST (Generative High-fidelity One Shot Transfer) is a one-shot face swap pipeline for images and videos, published as an IEEE paper and i… | 26 | 1582 | active |
| meta-pytorch/torchtune Torchtune is a PyTorch-native library for authoring, post-training, and experimenting with large language models. It provides hackable trai… | 69 | 5801 | maintenance |
| AlibabaResearch/DAMO-ConvAI The official codebase for Alibaba DAMO Academy's Conversational AI research, containing implementations of models from their published pape… | 71 | 1580 | active |
| SalesforceAIResearch/uni2ts Uni2TS is a PyTorch library for unified pre-training, fine-tuning, inference, and evaluation of universal time series forecasting transform… | 60 | 1580 | active |
| RWKV/rwkv.cpp A C++ port of the RWKV language model to the ggml tensor library, providing FP32, FP16, and quantized INT4/INT5/INT8 inference focused on C… | 37 | 1579 | active |
| tensorflow/model-optimization The TensorFlow Model Optimization Toolkit (tfmot) is a Python library providing tools to optimize machine learning models for deployment, i… | 82 | 1578 | stable |
| menyifang/MIMO MIMO is the official PyTorch implementation of a CVPR 2025 paper on controllable character video synthesis using spatially decomposed model… | 34 | 1578 | active |
| DiffEqML/torchdyn Torchdyn is a PyTorch library dedicated to numerical deep learning, providing tools for neural differential equations, implicit models, and… | 23 | 1578 | active |
| BloodAxe/pytorch-toolbelt A Python library of PyTorch extensions providing building blocks for fast R&D prototyping, including encoder-decoder architectures, special… | 44 | 1574 | active |
| FluxML/Zygote.jl Zygote.jl is a source-to-source automatic differentiation library for Julia that hooks into the Julia compiler to generate gradient (backwa… | 97 | 1568 | active |
| Xilinx/brevitas Brevitas is a PyTorch library for neural network quantization supporting both post-training quantization (PTQ) and quantization-aware train… | 91 | 1567 | active |
| Chinese-Text-Classification-Pytorch A PyTorch-based collection of ready-to-run Chinese text classification models including TextCNN, TextRNN, FastText, TextRCNN, BiLSTM with a… | 32 | 5728 | maintenance |
| ZeyueT/AudioX AudioX is a unified multimodal framework for anything-to-audio generation, producing audio and music conditioned on text, video, image, or … | 52 | 1552 | active |
| kijai/ComfyUI-CogVideoXWrapper A ComfyUI custom node wrapper for CogVideoX and related video generation models (including Fun variants, CogVideoX 1.5, and Go-with-the-Flo… | 39 | 1549 | active |
| Tencent/AngelSlim AngelSlim is a Python toolkit from Tencent for compressing large language models and related architectures (VLMs, diffusion, audio models) … | 72 | 1547 | active |
| mosaicml/streaming StreamingDataset is a Python library from MosaicML for fast, accurate streaming of training data from cloud object storage (S3, GCS, Azure,… | 70 | 1547 | active |
| lucidrains/soundstorm-pytorch A PyTorch implementation of SoundStorm, Google DeepMind's efficient parallel audio generation model that applies MaskGiT-style masked gener… | 39 | 1546 | active |
| facebookresearch/mmf MMF is a modular PyTorch framework for vision and language multimodal research from Facebook AI Research. It ships reference implementation… | 64 | 5633 | maintenance |
| google-ai-edge/model-explorer Model Explorer is a model graph visualizer and debugger from Google AI Edge that renders neural network graphs hierarchically with expandab… | 93 | 1542 | active |
| vibevoice-community/VibeVoice VibeVoice is a community-maintained fork of Microsoft's long-form conversational text-to-speech model, generating expressive multi-speaker … | 60 | 1542 | active |
| hustvl/MapTR MapTR is an end-to-end transformer-based framework for online vectorized HD map construction from camera imagery in autonomous driving. It … | 27 | 1542 | active |
| tensorflow/gnn TensorFlow GNN is a Python library for building Graph Neural Networks on TensorFlow, including a GraphTensor type for heterogeneous graphs,… | 66 | 1541 | active |
| lucidrains/DALLE-pytorch A PyTorch implementation/replication of OpenAI's DALL-E, a text-to-image transformer, including a discrete VAE and optional CLIP for rankin… | 23 | 5627 | maintenance |
| MoonshotAI/Moonlight Moonlight is a 3B/16B Mixture-of-Experts LLM trained with the Muon optimizer, released by Moonshot AI along with a memory- and communicatio… | 36 | 1540 | active |
| Arthur151/ROMP ROMP is a PyTorch-based library and pip-installable API (simple-romp) for real-time monocular multi-person 3D human mesh recovery, implemen… | 23 | 1538 | stable |
| xLLM-AI/xllm xLLM is a high-performance C++ inference engine for LLM, VLM, DiT and recommendation models, optimized for heterogeneous AI accelerators su… | 78 | 1536 | active |
| leela-zero/leela-zero Leela Zero is an open-source Go engine that reimplements AlphaGo Zero, combining Monte Carlo Tree Search with a deep residual convolutional… | 23 | 5586 | maintenance |
| JingyunLiang/SwinIR Official PyTorch implementation of SwinIR, a Swin Transformer-based model for image restoration tasks including super-resolution, denoising… | 23 | 5580 | maintenance |
| hustvl/LightningDiT LightningDiT is a research codebase for latent diffusion models implementing VA-VAE and LightningDiT, achieving FID 1.35 on ImageNet-256 wi… | 47 | 1529 | active |
| cchen156/Learning-to-See-in-the-Dark TensorFlow implementation of 'Learning to See in the Dark' (CVPR 2018), a deep learning model that brightens very dark, short-exposure RAW … | 51 | 5565 | maintenance |
| ByteDance-Seed/Triton-distributed Triton-distributed is a distributed compiler built on OpenAI Triton for computation-communication overlapping on multi-GPU systems. It lets… | 63 | 1526 | active |
| WenmuZhou/PytorchOCR A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP… | 59 | 1523 | active |
| keras-rl/keras-rl A Python library implementing state-of-the-art deep reinforcement learning algorithms (DQN, DDPG, SARSA, and more) that integrates seamless… | 23 | 5547 | maintenance |
| NirAharon/BoT-SORT BoT-SORT is a state-of-the-art multi-object tracker that combines motion and appearance information with camera motion compensation and an … | 32 | 1522 | active |
| caiyuanhao1998/Retinexformer Retinexformer is a one-stage Retinex-based Transformer model and toolbox for low-light image enhancement, published at ICCV 2023. It suppor… | 66 | 1518 | active |
| allenzren/open-pi-zero An open-source re-implementation of the pi0 vision-language-action (VLA) model from Physical Intelligence, built on a pre-trained PaliGemma… | 25 | 1518 | active |
| Phantom-video/Phantom Phantom is a subject-consistent video generation model from ByteDance that preserves reference subject identity via cross-modal alignment. … | 39 | 1517 | active |
| microsoft/Mage Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage… | 57 | 1516 | active |
| ZFTurbo/Music-Source-Separation-Training A Python training framework for music source separation models, supporting many architectures such as MDX23C, Demucs, Band Split RoFormer, … | 83 | 1512 | active |
| decoderesearch/SAELens SAELens is a Python library for training sparse autoencoders (SAEs) on language model activations and analyzing them for mechanistic interp… | 88 | 1510 | active |
| HiDream-ai/HiDream-O1-Image HiDream-O1-Image is an open-weights 8B image generation foundation model built on a Pixel-level Unified Transformer (UiT) that natively enc… | 53 | 1510 | active |
| imoneoi/openchat OpenChat is a library of open-source large language models fine-tuned with C-RLFT, an offline reinforcement learning strategy that learns f… | 20 | 5490 | maintenance |
| wb14123/seq2seq-couplet A deep learning project that generates Chinese couplets (对联) using a seq2seq model built with TensorFlow. It includes training scripts, a w… | 32 | 5487 | maintenance |
| Flashlight wav2letter++ is Facebook AI Research's end-to-end automatic speech recognition (ASR) toolkit written in C++. It has been consolidated into … | 62 | 5466 | maintenance |
| czczup/ViT-Adapter Official PyTorch implementation of ViT-Adapter, an ICLR 2023 Spotlight paper introducing a pre-training-free adapter that lets plain Vision… | 34 | 1503 | stable |
| charlesq34/pointnet Reference implementation of PointNet, a neural network architecture that directly consumes unordered 3D point clouds for classification and… | 32 | 5459 | maintenance |
| TencentARC/MotionCtrl MotionCtrl is the official implementation of a SIGGRAPH 2024 paper providing a unified and flexible motion controller for video generation … | 30 | 1500 | active |
| mattmireles/gemma-tuner-multimodal A Python tool for LoRA fine-tuning of Gemma 4 and 3n models on text, images, and audio using Apple Silicon's Metal Performance Shaders. It … | 62 | 1495 | active |
| ibab/tensorflow-wavenet A TensorFlow implementation of DeepMind's WaveNet generative neural network architecture for raw audio waveform generation. It provides tra… | 32 | 5428 | maintenance |
| open-mmlab/mmengine MMEngine is the foundational training engine library for OpenMMLab projects, providing a unified training loop, config system, registry, ho… | 71 | 1492 | active |
| bojone/bert4keras A lightweight, clean reimplementation of BERT and other transformer models (RoBERTa, ALBERT, T5, GPT, ELECTRA, NEZHA) for Keras/tf.keras. I… | 23 | 5415 | maintenance |
| google-deepmind/graph_nets DeepMind's library for building graph networks (graph neural networks) in TensorFlow and Sonnet, based on the 'Relational inductive biases,… | 32 | 5406 | maintenance |
| graphdeco-inria/diff-gaussian-rasterization A CUDA-based differentiable rasterization engine for 3D Gaussian Splatting, used in the SIGGRAPH 2023 paper '3D Gaussian Splatting for Real… | 29 | 1489 | stable |
| CUT3R/CUT3R CUT3R is the official PyTorch implementation of 'Continuous 3D Perception Model with Persistent State' (CVPR 2025 Oral), a stateful recurre… | 38 | 1486 | active |
| k2-fsa/icefall Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi… | 64 | 1482 | active |
| hustvl/DiffusionDrive DiffusionDrive is a truncated diffusion model for real-time end-to-end autonomous driving, released as the official PyTorch implementation … | 44 | 1480 | active |
| Soul-AILab/SoulX-FlashTalk SoulX-FlashTalk is a 14B audio-driven talking avatar model that streams infinite real-time video from a reference image and audio, achievin… | 58 | 1479 | active |
| alxndrTL/mamba.py A simple, readable pure-PyTorch (plus MLX) implementation of the Mamba state-space model architecture with a parallel scan for efficient tr… | 53 | 1478 | active |
| Om-Alve/smolGPT A minimal pure-PyTorch implementation for training small GPT-style LLMs from scratch, featuring flash attention, RMSNorm, SwiGLU, RoPE, and… | 23 | 1475 | active |
| mit-han-lab/torchsparse TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a… | 26 | 1472 | active |
| sczhou/Upscale-A-Video Upscale-A-Video is a diffusion-based model for real-world video super-resolution that takes low-resolution videos and text prompts as input… | 26 | 1470 | active |
| Lightning-AI/lightning-thunder Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a … | 72 | 1469 | active |
| autonomousvision/mip-splatting Mip-Splatting is a research implementation of alias-free 3D Gaussian Splatting, introducing a 3D smoothing filter and 2D Mip filter to elim… | 27 | 1466 | active |
| hojonathanho/diffusion The official reference implementation of Denoising Diffusion Probabilistic Models (DDPM) from the 2020 paper by Jonathan Ho et al., written… | 32 | 5300 | maintenance |
| NVIDIA/tacotron2 NVIDIA's PyTorch implementation of the Tacotron 2 text-to-speech model, which synthesizes mel spectrograms from text for vocoder-based audi… | 32 | 5296 | maintenance |
| yunjey/stargan Official PyTorch implementation of StarGAN, a unified generative adversarial network for multi-domain image-to-image translation (CVPR 2018… | 32 | 5296 | maintenance |
| FeiYull/TensorRT-Alpha A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl… | 32 | 1460 | active |
| tensorflow/tpu A collection of reference models and tools for training machine learning models on Google Cloud TPUs, maintained as a public mirror by the … | 72 | 5278 | maintenance |
| dbolya/yolact YOLACT is a PyTorch implementation of a fully convolutional model for real-time instance segmentation, accompanying the YOLACT and YOLACT++… | 51 | 5241 | maintenance |
| zylo117/Yet-Another-EfficientDet-Pytorch A PyTorch re-implementation of Google's EfficientDet object detection model with pretrained weights matching near-SOTA accuracy at real-tim… | 23 | 5238 | maintenance |
| amdegroot/ssd.pytorch A PyTorch implementation of the Single Shot MultiBox Detector (SSD) object detection model from the 2016 paper by Wei Liu et al. It include… | 32 | 5221 | maintenance |
| zsyOAOA/InvSR InvSR is a Python research library implementing arbitrary-steps image super-resolution via diffusion inversion, leveraging pre-trained diff… | 51 | 1443 | active |
| voice-cloning-app/Voice-Cloning-App A Python/PyTorch desktop application for cloning and synthesizing human voices from audio datasets. It handles the full pipeline from autom… | 23 | 1440 | active |
| google-deepmind/rlax RLax is a JAX-based library of building blocks for implementing reinforcement learning agents, providing mathematical operations like value… | 83 | 1439 | active |
| Walter0807/MotionBERT Official PyTorch implementation of MotionBERT (ICCV 2023), a unified pretrained model for learning human motion representations from 2D ske… | 65 | 1439 | active |
| tensorspace-team/tensorspace TensorSpace is a neural network 3D visualization framework built on TensorFlow.js, Three.js, and Tween.js. It provides Keras-like APIs to b… | 23 | 5191 | maintenance |
| tianweiy/DMD2 DMD2 is the official PyTorch implementation of Improved Distribution Matching Distillation, a NeurIPS 2024 method that distills diffusion m… | 28 | 1438 | active |
| neuralchen/SimSwap SimSwap is a PyTorch-based face-swapping framework that performs arbitrary face swaps on images and videos using a single trained model. It… | 23 | 5188 | maintenance |
| sql-machine-learning/sqlflow SQLFlow is a compiler that extends SQL with AI-oriented syntax (training, prediction, evaluation, explanation, and mathematical programming… | 23 | 5188 | maintenance |