function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ultralytics/yolov5 Ultralytics YOLOv5 is a PyTorch-based computer vision model family for real-time object detection, instance segmentation, and image classif… | 67 | 57929 | maintenance |
| thu-ml/tianshou Tianshou is a modular, high-performance deep reinforcement learning library built on pure PyTorch and Gymnasium. It offers both low-level h… | 72 | 10943 | active |
| triton-inference-server/server NVIDIA Triton Inference Server is an open-source inference serving software that deploys AI models from multiple frameworks (TensorRT, PyTo… | 98 | 10939 | stable |
| Lightricks/LTX-Video Official repository for LTX-Video, a DiT-based open-weights video generation model from Lightricks that generates high-fidelity video (up t… | 48 | 10907 | active |
| microsoft/TRELLIS.2 TRELLIS.2 is a 4B-parameter state-of-the-art 3D generative model from Microsoft for high-fidelity image-to-3D asset generation. It uses a f… | 57 | 10869 | active |
| cumulo-autumn/StreamDiffusion StreamDiffusion is a Python pipeline for real-time interactive diffusion-based image generation, achieving 100+ fps on modern GPUs. It opti… | 17 | 10806 | active |
| lucidrains/denoising-diffusion-pytorch A PyTorch implementation of Denoising Diffusion Probabilistic Models (DDPM) for generative image modeling, published on PyPI as denoising_d… | 87 | 10679 | active |
| autogluon/autogluon AutoGluon is an AutoML library that automates machine learning on tabular data, time series, text, and images with just a few lines of Pyth… | 90 | 10617 | active |
| Megvii-BaseDetection/YOLOX YOLOX is a high-performance anchor-free YOLO object detection model family implemented in PyTorch, with pretrained weights and export suppo… | 34 | 10587 | stable |
| facebookresearch/xformers xFormers is a PyTorch-based library of hackable, optimized Transformer building blocks with custom CUDA kernels for fast, memory-efficient … | 88 | 10542 | active |
| NVIDIA/cutlass CUTLASS is NVIDIA's collection of CUDA C++ template abstractions and Python DSLs for implementing high-performance GEMM and related linear … | 99 | 10317 | active |
| xai-org/grok-1 xAI's open release of the Grok-1 open-weights model (314B-parameter Mixture-of-Experts LLM) with JAX example code for loading and running i… | 25 | 52189 | maintenance |
| OpenGVLab/InternVL InternVL is a family of open-source multimodal large language models (vision-language models) that combine vision transformers with LLMs to… | 37 | 10146 | active |
| OpenMined/PySyft PySyft is a Python library that lets data scientists run computations on private data that stays on the data owner's server, with results s… | 74 | 9957 | active |
| facebookresearch/pytorch3d PyTorch3D is Facebook AI Research's library of efficient, reusable components for deep learning with 3D data, built on PyTorch. It provides… | 74 | 9954 | active |
| open-mmlab/mmsegmentation MMSegmentation is a PyTorch-based toolbox and benchmark for semantic segmentation, part of the OpenMMLab ecosystem. It provides implementat… | 23 | 9930 | stable |
| arogozhnikov/einops einops is a Python library providing readable, framework-agnostic tensor operations via mini-language functions like rearrange, reduce, and… | 77 | 9581 | stable |
| WongKinYiu/yolov9 Official PyTorch implementation of the YOLOv9 object detection paper, featuring Programmable Gradient Information for improved accuracy. It… | 16 | 9551 | active |
| modelscope/facechain FaceChain is a deep-learning toolchain from ModelScope for generating identity-preserved personal portraits (digital twins) from a single p… | 30 | 9508 | active |
| Oneflow-Inc/oneflow OneFlow is an open-source deep learning framework written in C++ with a PyTorch-like Python API, focused on scalable and efficient distribu… | 48 | 9428 | active |
| keras-team/autokeras AutoKeras is an AutoML library for deep learning built on Keras, developed by DATA Lab at Texas A&M University. It automates model architec… | 52 | 9326 | active |
| bytedance/monolith Monolith is a deep learning framework built on TensorFlow for large-scale recommendation modeling. It provides collisionless embedding tabl… | 10 | 9298 | active |
| LTX-2 Official Python package from Lightricks providing inference pipelines and LoRA training for LTX-2/LTX-2.5, an open-weights DiT-based founda… | 83 | 9260 | active |
| coqui-ai/TTS Coqui TTS is a deep learning toolkit for text-to-speech synthesis, providing pretrained models in over 1100 languages plus tools for traini… | 23 | 45953 | maintenance |
| pyro-ppl/pyro Pyro is a deep universal probabilistic programming library built on Python and PyTorch, supporting Bayesian modeling with variational infer… | 66 | 9037 | stable |
| dusty-nv/jetson-inference A C++/Python DNN inference library and tutorial guide ('Hello AI World') for deploying deep learning vision models on NVIDIA Jetson devices… | 44 | 8969 | stable |
| sebastianstarke/AI4Animation AI4Animation is a deep learning framework for data-driven character animation and control, built around Unity with a Python remake (AI4Anim… | 67 | 8846 | active |
| apple/ml-sharp SHARP is a Python tool from Apple that synthesizes a photorealistic 3D Gaussian splat representation from a single photograph in under a se… | 42 | 8843 | active |
| NVlabs/Sana SANA is an efficiency-oriented PyTorch codebase for high-resolution text-to-image and text-to-video generation built on Linear Diffusion Tr… | 74 | 8833 | active |
| MIC-DKFZ/nnUNet nnU-Net is a self-configuring deep learning framework for semantic image segmentation that automatically adapts preprocessing, U-Net archit… | 65 | 8829 | stable |
| FoundationVision/VAR Official PyTorch implementation of Visual Autoregressive Modeling (VAR), a NeurIPS 2024 Best Paper-winning method for scalable image genera… | 48 | 8729 | active |
| DepthAnything/Depth-Anything-V2 Depth Anything V2 is a foundation model for monocular depth estimation, trained on 595K synthetic labeled images and 62M+ real unlabeled im… | 56 | 8709 | stable |
| fudan-generative-vision/hallo Hallo is a Python research library implementing hierarchical audio-driven visual synthesis for animating portrait images into talking-head … | 14 | 8664 | active |
| MONAI MONAI is a PyTorch-based open-source framework for deep learning in healthcare imaging, providing domain-specific transforms, 3D architectu… | 87 | 8634 | stable |
| google-deepmind/alphafold3 Google DeepMind's official implementation of the AlphaFold 3 inference pipeline for predicting biomolecular structures and interactions. It… | 83 | 8495 | active |
| lucidrains/imagen-pytorch A PyTorch implementation of Imagen, Google's text-to-image neural network based on cascading DDPMs conditioned on T5 text embeddings. It pr… | 23 | 8424 | active |
| nl8590687/ASRT_SpeechRecognition ASRT is a deep-learning-based Chinese speech recognition (speech-to-text) system built with TensorFlow/Keras, using CNN, LSTM, attention me… | 57 | 8383 | active |
| XPixelGroup/BasicSR BasicSR is an open-source PyTorch toolbox for image and video restoration tasks such as super-resolution, denoising, deblurring, and JPEG a… | 23 | 8367 | stable |
| QwenLM/Qwen-Image Qwen-Image is a 20B MMDiT image generation foundation model from the Qwen team, with strong complex text rendering (especially Chinese) and… | 48 | 8265 | active |
| Ucas-HaoranWei/GOT-OCR2.0 Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, … | 25 | 8216 | active |
| google-research/bert Google Research's official TensorFlow implementation of BERT, the Bidirectional Encoder Representations from Transformers language model, a… | 10 | 40046 | maintenance |
| shenweichen/DeepCTR DeepCTR is a Python library of easy-to-use, modular, and extendible deep-learning based CTR (click-through rate) prediction models built on… | 77 | 8050 | stable |
| NVIDIA/Isaac-GR00T NVIDIA Isaac GR00T N1.7 is an open vision-language-action (VLA) foundation model for generalized humanoid robot skills, taking language and… | 76 | 7926 | stable |
| stanfordnlp/stanza Stanza is the Stanford NLP Group's official Python library for linguistic analysis of human language text. It provides a neural pipeline bu… | 97 | 7867 | stable |
| wang-xinyu/tensorrtx A C++ collection of popular deep learning networks (YOLO variants, ResNet, MobileNet, Swin Transformer, OCR models, and more) implemented f… | 76 | 7827 | active |
| TencentARC/GFPGAN GFPGAN is a Python library built on PyTorch that restores and enhances real-world degraded face photos using GAN-based priors. It provides … | 23 | 37657 | maintenance |
| PaddlePaddle/ERNIE Official repository for Baidu's ERNIE 4.5 family of large multimodal models and ERNIEKit, an industrial-grade training toolkit built on Pad… | 65 | 7738 | active |
| MeiGen-AI/InfiniteTalk InfiniteTalk is an open-source model and framework for unlimited-length audio-driven talking video generation, supporting both image-to-vid… | 55 | 7700 | active |
| meituan-longcat/LongCat-Video LongCat-Video is a 13.6B-parameter foundational video generation model from Meituan that unifies text-to-video, image-to-video, and video-c… | 54 | 7665 | active |
| google-deepmind/weathernext Google DeepMind's WeatherNext family of AI weather forecasting models, including WeatherNext 2, GraphCast, and GenCast, with code and pretr… | 82 | 7596 | active |
| h2oai/h2o-3 H2O-3 is an open-source, distributed, in-memory machine learning platform implementing algorithms such as GLM, GBM/XGBoost, Random Forest, … | 77 | 7494 | stable |
| SkyworkAI/SkyReels-V2 SkyReels-V2 is an open-source infinite-length film/video generative model using an AutoRegressive Diffusion-Forcing architecture, released … | 48 | 7462 | active |
| EleutherAI/gpt-neox GPT-NeoX is EleutherAI's library for training large-scale autoregressive transformer language models on GPUs, built on NVIDIA's Megatron an… | 62 | 7459 | active |
| LargeWorldModel/LWM Large World Model (LWM) is a family of open-source 7B-parameter multimodal autoregressive transformer models trained on long videos and boo… | 25 | 7425 | active |
| apple/ml-fastvlm Official implementation of FastVLM, a vision language model with an efficient hybrid vision encoder (FastViTHD) that reduces token count an… | 28 | 7411 | active |
| facebookresearch/SlowFast PySlowFast is a PyTorch-based open-source video understanding codebase from Facebook AI Research (FAIR). It provides implementations of sta… | 65 | 7410 | active |
| facebookresearch/sam-3d-objects SAM 3D Objects is a foundation model from Meta that reconstructs full 3D shape geometry, texture, and layout from a single image, with code… | 55 | 7322 | active |
| google/flax Flax is a neural network library and ecosystem for JAX designed for flexibility, featuring the newer NNX API with first-class Python refere… | 99 | 7303 | active |
| PaddlePaddle/Paddle-Lite Paddle Lite is a high-performance, lightweight deep learning inference engine from Baidu's PaddlePaddle ecosystem, designed for mobile, emb… | 58 | 7273 | active |
| mit-han-lab/streaming-llm StreamingLLM is a research framework from MIT Han Lab implementing the Attention Sinks method (ICLR 2024) for efficient streaming language … | 27 | 7268 | stable |
| BVLC/caffe Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp… | 23 | 34556 | maintenance |
| CMU-Perceptual-Computing-Lab/openpose OpenPose is a real-time multi-person keypoint detection library that jointly estimates human body, face, hand, and foot keypoints (135 tota… | 23 | 34413 | maintenance |
| ControlNet ControlNet is a neural network architecture that adds conditional control (edges, poses, depth, etc.) to pretrained text-to-image diffusion… | 31 | 34091 | maintenance |
| flwrlabs/flower Flower (flwr) is an open-source Python framework for building federated and collaborative AI systems, supporting any ML framework such as P… | 99 | 7085 | active |
| threestudio-project/threestudio threestudio is a unified open-source framework for 3D content generation from text prompts, single images, and few-shot images by lifting 2… | 20 | 7059 | active |
| zai-org/GLM-5 GLM-5 is a family of open-weight large language models from Z.ai (Zhipu AI), including the flagship GLM-5.2 with 1M-token context and the m… | 60 | 7057 | active |
| SakanaAI/AI-Scientist-v2 An end-to-end autonomous AI research system that generates hypotheses, runs machine learning experiments, analyzes data, and writes complet… | 46 | 7050 | active |
| google/gemma.cpp A lightweight, standalone C++ inference engine for Google's Gemma foundation models (Gemma 2/3, PaliGemma 2), with a small ~2K LoC core and… | 72 | 7031 | active |
| deepseek-ai/DeepSeek-Coder-V2 DeepSeek-Coder-V2 is an open-weight Mixture-of-Experts code language model family released by DeepSeek AI, with weights available on Huggin… | 47 | 7007 | active |
| apple/corenet CoreNet is Apple's deep neural network training toolkit for training standard and novel small and large-scale models, including foundation … | 45 | 7007 | active |
| deeppavlov/DeepPavlov DeepPavlov is an open-source Python NLP library built on PyTorch and Hugging Face transformers for developing, training, and deploying stat… | 39 | 6989 | active |
| deepchem/deepchem DeepChem is a Python library that democratizes deep learning for drug discovery, quantum chemistry, materials science, and biology. It prov… | 67 | 6961 | active |
| PaddlePaddle/models PaddlePaddle's officially maintained industry-grade model repository containing 600+ models across computer vision, NLP, speech, recommenda… | 23 | 6932 | active |
| VAST-AI-Research/TripoSR TripoSR is an open-source model for fast feedforward 3D object reconstruction from a single image, developed by Tripo AI and Stability AI. … | 64 | 6888 | active |
| eriklindernoren/ML-From-Scratch A Python library providing bare-bones NumPy implementations of fundamental machine learning models and algorithms, from linear regression t… | 32 | 32528 | maintenance |
| facebookresearch/fairseq Fairseq is a PyTorch-based sequence modeling toolkit from Facebook AI Research for training custom models for translation, summarization, l… | 10 | 32231 | maintenance |
| tencent-ailab/IP-Adapter IP-Adapter is a lightweight (22M parameter) adapter that adds image prompt capability to pretrained text-to-image diffusion models like Sta… | 28 | 6677 | stable |
| yangjianxin1/Firefly Firefly is an open-source one-stop training tool for large language models, supporting pretraining, instruction fine-tuning (SFT), and DPO … | 21 | 6653 | active |
| ml5js/ml5-library ml5.js is a friendly, beginner-oriented JavaScript machine learning library for the browser, built on top of TensorFlow.js. It provides acc… | 23 | 6587 | active |
| FareedKhan-dev/kimi-k3-in-c A dependency-free C99 inference engine that runs the 2.78-trillion-parameter Kimi K3 model on a single CPU with as little as 8 GB of RAM by… | 79 | 6524 | active |
| open-mmlab/mmdetection3d MMDetection3D is OpenMMLab's next-generation platform for general 3D object detection, built on PyTorch. It provides a modular toolbox with… | 23 | 6518 | active |
| open-mmlab/mmcv MMCV is the foundational computer vision library for the OpenMMLab ecosystem, providing image/video I/O, data transformations, and CUDA ope… | 52 | 6470 | stable |
| TMElyralab/MuseTalk MuseTalk is a real-time, high-fidelity lip-sync model that modifies a face region in video according to input audio via latent space inpain… | 45 | 6459 | active |
| haifengl/smile SMILE is a comprehensive, high-performance machine learning framework for the JVM with idiomatic APIs for Java, Scala, and Kotlin. It cover… | 99 | 6413 | active |
| KevinMusgrave/pytorch-metric-learning A PyTorch library providing modular losses, miners, samplers, and testers for deep metric learning. It supports building complete train/tes… | 48 | 6339 | active |
| yl4579/StyleTTS2 StyleTTS 2 is a PyTorch text-to-speech model that uses style diffusion and adversarial training with large speech language models (e.g., Wa… | 29 | 6336 | active |
| mindee/doctr docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe… | 90 | 6315 | active |
| RangiLyu/nanodet NanoDet-Plus is a super fast, lightweight anchor-free object detection model implemented in PyTorch, with model sizes as small as 980KB (IN… | 23 | 6252 | stable |
| ByteDance-Seed/Depth-Anything-3 Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,… | 59 | 6213 | active |
| skorch-dev/skorch skorch is a Python library that wraps PyTorch neural networks in a scikit-learn compatible API, providing estimators like NeuralNetClassifi… | 85 | 6173 | active |
| ByteDance-Seed/Bagel BAGEL is an open-source unified multimodal foundation model with 7B active parameters (14B total) trained on interleaved multimodal data. I… | 55 | 6159 | active |
| timeseriesAI/tsai tsai is an open-source deep learning library built on PyTorch and fastai for time series and sequential data tasks such as classification, … | 84 | 6111 | active |
| deezer/spleeter Spleeter is Deezer's music source separation library with pretrained TensorFlow models that splits audio into stems (vocals, drums, bass, p… | 62 | 28402 | maintenance |
| open-edge-platform/anomalib Anomalib is a Python deep learning library for anomaly detection, offering state-of-the-art unsupervised algorithms for detecting and local… | 98 | 6088 | active |
| svc-develop-team/so-vits-svc A deep learning framework based on SoftVC VITS for singing voice conversion (SVC), letting users train models that convert one singing voic… | 10 | 28125 | maintenance |
| bytedance/LatentSync LatentSync is an end-to-end lip-sync framework from ByteDance based on audio-conditioned latent diffusion models, using Stable Diffusion to… | 33 | 6026 | active |
| Doubiiu/ToonCrafter ToonCrafter is a generative model that interpolates two cartoon images into a short animation by leveraging pre-trained image-to-video diff… | 29 | 6003 | stable |
| OFA-Sys/Chinese-CLIP Chinese-CLIP is a Chinese version of the CLIP model trained on ~200 million Chinese image-text pairs, built on open_clip. It provides APIs,… | 66 | 5998 | active |
| z-lab/dflash DFlash is a lightweight block diffusion model used as a draft model for speculative decoding of large language models, drafting entire toke… | 71 | 5967 | active |
| lucidrains/x-transformers A concise PyTorch library implementing full-attention transformer architectures (encoder, decoder, encoder-decoder, and vision transformers… | 85 | 5942 | active |