domain: deep-learning
2771 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| wang-xinyu/tensorrtx A C++ collection of popular deep learning networks (YOLO variants, ResNet, MobileNet, Swin Transformer, OCR models, and more) implemented f… | 76 | 7827 | active |
| TencentARC/GFPGAN GFPGAN is a Python library built on PyTorch that restores and enhances real-world degraded face photos using GAN-based priors. It provides … | 23 | 37657 | maintenance |
| deepseek-ai/DeepGEMM DeepGEMM is a high-performance CUDA BLAS kernel library for NVIDIA tensor cores, providing FP8, FP4, and BF16 GEMMs plus fused MoE and othe… | 84 | 7738 | active |
| PaddlePaddle/ERNIE Official repository for Baidu's ERNIE 4.5 family of large multimodal models and ERNIEKit, an industrial-grade training toolkit built on Pad… | 65 | 7738 | active |
| MeiGen-AI/InfiniteTalk InfiniteTalk is an open-source model and framework for unlimited-length audio-driven talking video generation, supporting both image-to-vid… | 55 | 7700 | active |
| meta-llama/llama-models Meta's official repository of Llama large language models with Python utilities for working with them, including model cards, licenses, and… | 55 | 7685 | active |
| meituan-longcat/LongCat-Video LongCat-Video is a 13.6B-parameter foundational video generation model from Meituan that unifies text-to-video, image-to-video, and video-c… | 54 | 7665 | active |
| babysor/MockingBird MockingBird is a PyTorch-based AI voice cloning toolbox that can clone a voice from a 5-second sample and generate arbitrary speech in real… | 54 | 36909 | maintenance |
| 1adrianb/face-alignment A Python library built on PyTorch that detects 2D and 3D facial landmarks in images using the FAN deep learning face alignment network. It … | 70 | 7538 | active |
| SkyworkAI/SkyReels-V2 SkyReels-V2 is an open-source infinite-length film/video generative model using an AutoRegressive Diffusion-Forcing architecture, released … | 48 | 7462 | active |
| EleutherAI/gpt-neox GPT-NeoX is EleutherAI's library for training large-scale autoregressive transformer language models on GPUs, built on NVIDIA's Megatron an… | 62 | 7459 | active |
| LargeWorldModel/LWM Large World Model (LWM) is a family of open-source 7B-parameter multimodal autoregressive transformer models trained on long videos and boo… | 25 | 7425 | active |
| apple/ml-fastvlm Official implementation of FastVLM, a vision language model with an efficient hybrid vision encoder (FastViTHD) that reduces token count an… | 28 | 7411 | active |
| facebookresearch/SlowFast PySlowFast is a PyTorch-based open-source video understanding codebase from Facebook AI Research (FAIR). It provides implementations of sta… | 65 | 7410 | active |
| facebookresearch/sam-3d-objects SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code… | 55 | 7322 | active |
| vladmandic/sdnext SD.Next is an open-source, self-hosted WebUI server application for AI generative image and video creation built on Stable Diffusion and Di… | 76 | 7320 | active |
| arcee-ai/mergekit mergekit is a Python toolkit for merging pre-trained large language models directly in weight space, supporting many merge methods (SLERP, … | 63 | 7310 | active |
| google/flax Flax is a neural network library and ecosystem for JAX designed for flexibility, featuring the newer NNX API with first-class Python refere… | 99 | 7303 | active |
| PaddlePaddle/Paddle-Lite Paddle Lite is a high-performance, lightweight deep learning inference engine from Baidu's PaddlePaddle ecosystem, designed for mobile, emb… | 58 | 7273 | active |
| mit-han-lab/streaming-llm StreamingLLM is a research framework from MIT Han Lab implementing the Attention Sinks method (ICLR 2024) for efficient streaming language … | 27 | 7268 | stable |
| Zyphra/Zonos Zonos-v0.1 is an open-weight text-to-speech model trained on over 200k hours of multilingual speech, with a Python library for inference. I… | 24 | 7244 | active |
| kohya-ss/sd-scripts A collection of Python training, generation, and utility scripts for Stable Diffusion and other image generation models, most widely used f… | 89 | 7210 | active |
| BVLC/caffe Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp… | 23 | 34556 | maintenance |
| tensorflow/tensorboard TensorBoard is TensorFlow's visualization toolkit, a suite of web applications for inspecting and understanding machine learning training r… | 88 | 7206 | active |
| CMU-Perceptual-Computing-Lab/openpose OpenPose is a real-time multi-person keypoint detection library that jointly estimates human body, face, hand, and foot keypoints (135 tota… | 23 | 34413 | maintenance |
| ControlNet ControlNet is a neural network architecture that adds conditional control (edges, poses, depth, etc.) to pretrained text-to-image diffusion… | 31 | 34091 | maintenance |
| yangchris11/samurai SAMURAI is the official implementation of a zero-shot visual object tracker built on top of Segment Anything Model 2 (SAM 2), using a motio… | 27 | 7112 | active |
| threestudio-project/threestudio threestudio is a unified open-source framework for 3D content generation from text prompts, single images, and few-shot images by lifting 2… | 20 | 7059 | active |
| deepseek-ai/DeepSpec DeepSpec is a full-stack Python codebase from DeepSeek for training and evaluating draft models used in speculative decoding of large langu… | 54 | 7041 | active |
| apple/corenet CoreNet is Apple's deep neural network training toolkit for training standard and novel small and large-scale models, including foundation … | 45 | 7007 | active |
| TheTom/turboquant_plus TurboQuant+ is a Python reference implementation of the TurboQuant KV cache compression method (ICLR 2026), using PolarQuant codebooks and … | 56 | 7006 | active |
| deeppavlov/DeepPavlov DeepPavlov is an open-source Python NLP library built on PyTorch and Hugging Face transformers for developing, training, and deploying stat… | 39 | 6989 | active |
| PaddlePaddle/models PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda… | 23 | 6932 | active |
| eriklindernoren/ML-From-Scratch A Python library providing bare-bones NumPy implementations of fundamental machine learning models and algorithms, from linear regression t… | 32 | 32528 | maintenance |
| clearml/clearml ClearML is an open-source Python SDK and platform for MLOps/LLMOps that provides auto-logged experiment tracking, data versioning, pipeline… | 99 | 6840 | stable |
| facebookresearch/fairseq Fairseq is a PyTorch-based sequence modeling toolkit from Facebook AI Research for training custom models for translation, summarization, l… | 10 | 32231 | maintenance |
| kijai/ComfyUI-WanVideoWrapper ComfyUI custom nodes wrapping WanVideo (Wan2.1) and related video generation models. It provides a standalone sandbox for quickly implement… | 58 | 6678 | active |
| tencent-ailab/IP-Adapter IP-Adapter is a lightweight (22M parameter) adapter that adds image prompt capability to pretrained text-to-image diffusion models like Sta… | 28 | 6677 | stable |
| FoundationVision/ByteTrack ByteTrack is a PyTorch-based multi-object tracking (MOT) library implementing the ECCV 2022 paper 'Multi-Object Tracking by Associating Eve… | 32 | 6654 | stable |
| yangjianxin1/Firefly Firefly is an open-source one-stop training tool for large language models, supporting pretraining, instruction fine-tuning (SFT), and DPO … | 21 | 6653 | active |
| AILab-CVC/YOLO-World YOLO-World is a real-time open-vocabulary object detection model and Python toolkit from Tencent AI Lab and HUST, published at CVPR 2024. I… | 29 | 6529 | active |
| FareedKhan-dev/kimi-k3-in-c A dependency-free C99 inference engine that runs the 2.78-trillion-parameter Kimi K3 model on a single CPU with as little as 8 GB of RAM by… | 79 | 6524 | active |
| open-mmlab/mmdetection3d MMDetection3D is OpenMMLab's next-generation platform for general 3D object detection, built on PyTorch. It provides a modular toolbox with… | 23 | 6518 | active |
| rtqichen/torchdiffeq torchdiffeq is a PyTorch library of differentiable ordinary differential equation (ODE) solvers, best known as the canonical implementation… | 39 | 6477 | stable |
| open-mmlab/mmcv MMCV is the foundational computer vision library for the OpenMMLab ecosystem, providing image/video I/O, data transformations, and CUDA ope… | 52 | 6470 | stable |
| TMElyralab/MuseTalk MuseTalk is a real-time, high-fidelity lip-sync model that modifies a face region in video according to input audio via latent space inpain… | 45 | 6459 | active |
| HVision-NKU/StoryDiffusion StoryDiffusion is the official implementation of a NeurIPS 2024 Spotlight paper introducing Consistent Self-Attention for character-consist… | 24 | 6452 | active |
| haifengl/smile SMILE is a comprehensive, high-performance machine learning framework for the JVM with idiomatic APIs for Java, Scala, and Kotlin. It cover… | 99 | 6413 | active |
| multimodal-art-projection/YuE YuE is a family of open-source foundation models based on the LLaMA2 architecture that generate full songs (up to five minutes) with vocals… | 32 | 6403 | active |
| tensorflow/serving TensorFlow Serving is a flexible, high-performance serving system for machine learning models designed for production environments. It mana… | 86 | 6360 | stable |
| KevinMusgrave/pytorch-metric-learning A PyTorch library providing modular losses, miners, samplers, and testers for deep metric learning. It supports building complete train/tes… | 48 | 6339 | active |
| yl4579/StyleTTS2 StyleTTS 2 is a PyTorch text-to-speech model that uses style diffusion and adversarial training with large speech language models (e.g., Wa… | 29 | 6336 | active |
| mindee/doctr docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe… | 90 | 6315 | active |
| canopyai/Orpheus-TTS Orpheus TTS is an open-source text-to-speech system built on a Llama-3b backbone that produces human-sounding speech with emotion control a… | 45 | 6314 | active |
| meta-llama/llama3 The official Meta repository for Llama 3, providing model weights download scripts, tokenizer, and minimal example code for running inferen… | 10 | 29247 | maintenance |
| flashinfer-ai/flashinfer FlashInfer is a GPU kernel library and kernel generator for LLM inference, providing unified APIs for attention, GEMM, and MoE operations w… | 90 | 6252 | active |
| RangiLyu/nanodet NanoDet-Plus is a super fast, lightweight anchor-free object detection model implemented in PyTorch, with model sizes as small as 980KB (IN… | 23 | 6252 | stable |
| PaddlePaddle/PaddleX PaddleX is a low-code, all-in-one AI development tool built on the PaddlePaddle framework, bundling 200+ pretrained models into 33 producti… | 92 | 6251 | active |
| meta-pytorch/gpt-fast A minimal (<1000 lines) PyTorch-native implementation of fast transformer text generation, demonstrating low-latency LLM inference with int… | 44 | 6249 | active |
| ByteDance-Seed/Depth-Anything-3 Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,… | 59 | 6213 | active |
| skorch-dev/skorch skorch is a Python library that wraps PyTorch neural networks in a scikit-learn compatible API, providing estimators like NeuralNetClassifi… | 85 | 6173 | active |
| timeseriesAI/tsai tsai is an open-source deep learning library built on PyTorch and fastai for time series and sequential data tasks such as classification, … | 84 | 6111 | active |
| Akegarasu/lora-scripts SD-Trainer is a GUI application and set of scripts for training LoRA and Dreambooth fine-tunes of Stable Diffusion diffusion models, wrappi… | 66 | 6110 | active |
| bytedance/MegaTTS3 MegaTTS 3 is ByteDance's open-source PyTorch text-to-speech model with a lightweight 0.45B-parameter Diffusion Transformer backbone. It pro… | 59 | 6091 | active |
| open-edge-platform/anomalib Anomalib is a Python deep learning library for anomaly detection, offering state-of-the-art unsupervised algorithms for detecting and local… | 98 | 6088 | active |
| svc-develop-team/so-vits-svc A deep learning framework based on SoftVC VITS for singing voice conversion (SVC), letting users train models that convert one singing voic… | 10 | 28125 | maintenance |
| bytedance/LatentSync LatentSync is an end-to-end lip-sync framework from ByteDance based on audio-conditioned latent diffusion models, using Stable Diffusion to… | 33 | 6026 | active |
| om-ai-lab/VLM-R1 VLM-R1 is a framework for training R1-style large vision-language models using reinforcement learning (GRPO) on top of Qwen2.5-VL. It provi… | 63 | 6015 | active |
| OFA-Sys/Chinese-CLIP Chinese-CLIP is a Chinese version of the CLIP model trained on ~200 million Chinese image-text pairs, built on open_clip. It provides APIs,… | 66 | 5998 | active |
| z-lab/dflash DFlash is a lightweight block diffusion model used as a draft model for speculative decoding of large language models, drafting entire toke… | 71 | 5967 | active |
| lucidrains/x-transformers A concise PyTorch library implementing full-attention transformer architectures (encoder, decoder, encoder-decoder, and vision transformers… | 85 | 5942 | active |
| pjreddie/darknet Darknet is an open-source neural network framework written in C and CUDA, best known as the original home of the YOLO real-time object dete… | 32 | 26492 | maintenance |
| DeepLabCut/DeepLabCut DeepLabCut is an open-source Python toolbox for markerless 2D and 3D pose estimation of user-defined body parts using deep neural networks … | 89 | 5745 | stable |
| NVIDIA/DALI NVIDIA DALI is a GPU-accelerated data loading and preprocessing library with optimized building blocks and an execution engine for deep lea… | 92 | 5734 | active |
| google/gemma_pytorch The official PyTorch implementation of Google's Gemma family of open large language models, including text-only and multimodal variants. It… | 10 | 5719 | active |
| google-deepmind/gemma The official JAX-based Python library from Google DeepMind for running, sampling from, and fine-tuning the Gemma family of open-weight larg… | 87 | 5695 | active |
| meta-pytorch/captum Captum is a model interpretability and understanding library for PyTorch, providing implementations of algorithms like Integrated Gradients… | 81 | 5693 | active |
| open-mmlab/OpenPCDet OpenPCDet is a PyTorch-based open-source toolbox for LiDAR-based 3D object detection. It provides official implementations of models like P… | 53 | 5692 | active |
| alexlenail/NN-SVG A web-based tool for creating publication-ready neural network architecture diagrams parametrically, supporting FCNN, LeNet-style CNN, and … | 61 | 5686 | stable |
| apple/ml-depth-pro Depth Pro is Apple's reference implementation of a foundation model for zero-shot metric monocular depth estimation, producing sharp high-r… | 30 | 5683 | active |
| huggingface/alignment-handbook A collection of robust training recipes and scripts from Hugging Face for aligning large language models with human and AI preferences, cov… | 66 | 5671 | active |
| OpenSenseNova/SenseNova-U1 SenseNova-U is a series of open-weight unified multimodal models (e.g., SenseNova-U1.5-8B-MoT) built on the NEO-unify architecture that com… | 59 | 5668 | active |
| pytorch/torchtitan torchtitan is a PyTorch-native platform for large-scale training of generative AI models, offering a clean-room implementation of PyTorch's… | 79 | 5667 | active |
| Fanghua-Yu/SUPIR SUPIR is a Python-based photo-realistic image restoration system built on SDXL diffusion priors and LLaVA captioning, presented at CVPR 202… | 36 | 5649 | active |
| fla-org/flash-linear-attention A PyTorch library providing hardware-efficient implementations of emerging sequence model architectures, including linear attention, sparse… | 88 | 5627 | active |
| huggingface/parler-tts Parler-TTS is a lightweight text-to-speech library from Hugging Face that generates high-quality, natural-sounding speech controllable via … | 25 | 5586 | active |
| matterport/Mask_RCNN A Python implementation of the Mask R-CNN model for object detection and instance segmentation, built on Keras and TensorFlow with a ResNet… | 23 | 25567 | maintenance |
| mosaicml/composer Composer is an open-source PyTorch-based deep learning training library by MosaicML (now Databricks) for training neural networks faster an… | 65 | 5495 | active |
| LaurentMazare/tch-rs tch-rs is a Rust crate providing thin bindings to the C++ API of PyTorch (libtorch), staying close to the original API. It enables tensor o… | 67 | 5479 | active |
| lyuwenyu/RT-DETR Official implementation of RT-DETR and RT-DETRv2, real-time object detection transformers that outperform YOLO models, in PyTorch and Paddl… | 74 | 5476 | active |
| tensorflow/rust TensorFlow Rust provides idiomatic Rust language bindings for TensorFlow via its C API. It lets Rust programs build and run TensorFlow comp… | 10 | 5476 | active |
| obss/sahi SAHI (Slicing Aided Hyper Inference) is a Python vision library for detecting small objects in large images via sliced/tiled inference, wor… | 99 | 5473 | active |
| HarisIqbal88/PlotNeuralNet A LaTeX/TikZ-based library for drawing neural network architecture diagrams, with a Python interface for generating the TikZ code. It is co… | 23 | 24955 | maintenance |
| karpathy/minGPT A minimal, clean PyTorch re-implementation of OpenAI's GPT covering both training and inference in roughly 300 lines of code. It is designe… | 32 | 24840 | maintenance |
| isl-org/MiDaS MiDaS is a Python library with pretrained models for robust monocular depth estimation from a single image, based on the TPAMI 2022 paper a… | 10 | 5420 | stable |
| facebookresearch/sapiens Sapiens is a family of foundation models from Meta Reality Labs for human-centric vision tasks including 2D pose estimation, body-part segm… | 61 | 5418 | active |
| apple/coremltools Apple's official Python package for converting machine learning models from TensorFlow, PyTorch, scikit-learn, XGBoost, and LibSVM into the… | 80 | 5399 | active |
| InternLM/xtuner XTuner is an open-source LLM training engine from InternLM designed for fine-tuning ultra-large-scale Mixture-of-Experts (MoE) models, with… | 67 | 5183 | active |
| transformerlab/transformerlab-app Transformer Lab is an open-source desktop application (built with Electron and Python) that provides a unified GUI for training, fine-tunin… | 84 | 5179 | active |
| h2oai/h2o-llmstudio H2O LLM Studio is a framework and no-code GUI for fine-tuning state-of-the-art large language models, built by H2O.ai. It supports LoRA and… | 96 | 5172 | active |