function: deep-learning
2653 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| AlmondGod/tinyworlds A minimal Python implementation of DeepMind's Genie autoregressive world model, including a video tokenizer, action tokenizer, and dynamics… | 54 | 1378 | active |
| keyu-tian/SparK SparK is the official PyTorch implementation of an ICLR 2023 Spotlight paper that applies BERT/MAE-style masked image modeling to convoluti… | 22 | 1376 | stable |
| qubvel/segmentation_models A Python library providing neural network architectures for image segmentation (Unet, FPN, Linknet, PSPNet) built on Keras and TensorFlow K… | 23 | 4923 | maintenance |
| minimaxir/textgenrnn A Python 3 library built on Keras/TensorFlow for easily training char-rnn style neural networks that generate text from any dataset in a fe… | 23 | 4922 | maintenance |
| Meituan-AutoML/MobileVLM MobileVLM is a family of compact vision language models (1.4B-3B parameters) designed to run efficiently on mobile devices, combining small… | 17 | 1370 | active |
| pnnl/neuromancer NeuroMANCER is a PyTorch-based differentiable programming library for solving parametric constrained optimization problems, physics-informe… | 72 | 1369 | active |
| OpenPPL/ppl.nn PPLNN is a high-performance deep-learning inference engine written in C++ that runs ONNX models on x86 CPUs and NVIDIA GPUs, with a dedicat… | 32 | 1367 | active |
| xandergos/terrain-diffusion Terrain Diffusion is a Python framework that uses diffusion models as a learned, deterministic replacement for Perlin noise, generating inf… | 64 | 1365 | active |
| ImprintLab/MedSegDiff MedSegDiff is a diffusion probabilistic model framework for segmenting and reconstructing organs and tissues from medical images, with a tr… | 50 | 1363 | active |
| hustvl/VAD VAD is an end-to-end autonomous driving framework that models the driving scene as a fully vectorized representation of agents and map elem… | 60 | 1362 | active |
| yuantianyuan01/FastWAM Official PyTorch codebase for Fast-WAM, a World Action Model for robot manipulation that skips test-time future video imagination and gener… | 59 | 1362 | active |
| bytedance/UNO UNO is a research framework from ByteDance for subject-driven image generation with diffusion transformers, supporting both single- and mul… | 38 | 1362 | active |
| MoonInTheRiver/DiffSinger Official PyTorch implementation of DiffSinger, an AAAI 2022 paper on singing voice synthesis and text-to-speech using a shallow diffusion m… | 65 | 4851 | maintenance |
| huggingface/finetrainers finetrainers is a Hugging Face library for scalable, memory-optimized training (fine-tuning) of diffusion models, including LoRA training o… | 62 | 1358 | active |
| Sense-X/Co-DETR Co-DETR is a PyTorch implementation of DETRs with Collaborative Hybrid Assignments Training, an ICCV 2023 object detection and instance seg… | 32 | 1357 | stable |
| blei-lab/edward Edward is a Python library for probabilistic modeling, inference, and criticism built on TensorFlow. It supports deep generative models, va… | 23 | 4843 | maintenance |
| reiniscimurs/DRL-robot-navigation A ROS Gazebo simulation project that trains a mobile robot to navigate to random goals while avoiding obstacles using a TD3 deep reinforcem… | 58 | 1356 | active |
| mega-sam/mega-sam MegaSaM is a research codebase implementing a deep visual SLAM system that estimates camera parameters and consistent depth maps from casua… | 48 | 1355 | active |
| k2-fsa/k2 k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)… | 64 | 1352 | active |
| shivammehta25/Matcha-TTS Matcha-TTS is a PyTorch-based text-to-speech system that uses conditional flow matching for fast, non-autoregressive speech synthesis. It s… | 62 | 1349 | active |
| wyhuai/DDNM DDNM is a Python research codebase implementing the Denoising Diffusion Null-Space Model for zero-shot image restoration, published as an I… | 32 | 1349 | stable |
| MegEngine/MegEngine MegEngine is a fast, scalable deep learning framework with automatic differentiation, developed in C++ with Python bindings. It unifies tra… | 23 | 4808 | maintenance |
| sjvasquez/handwriting-synthesis A Python implementation of Alex Graves' handwriting synthesis experiments using recurrent neural networks, generating realistic handwritten… | 32 | 4802 | maintenance |
| PKU-VCL-3DV/SLAM3R SLAM3R is a real-time dense 3D scene reconstruction system that regresses 3D points from monocular RGB video using feed-forward neural netw… | 42 | 1344 | active |
| muzishen/IMAGDressing IMAGDressing-v1 is a diffusion-based framework for customizable virtual dressing that generates human images with fixed garments and contro… | 44 | 1343 | active |
| FireRedTeam/FireRed-Image-Edit FireRed-Image-Edit is an open-source image editing foundation model built on diffusion models, released as PyTorch model weights with infer… | 49 | 1341 | active |
| alibaba/graph-learn Graph-Learn (formerly AliGraph) is a distributed framework for developing and applying large-scale graph neural networks, with a training l… | 36 | 1341 | active |
| rstudio/tensorflow An R package that provides full access to the TensorFlow API from R via reticulate, bridging R users to TensorFlow's Python implementation.… | 61 | 1339 | active |
| PKU-YuanGroup/MagicTime MagicTime is a metamorphic time-lapse video generation pipeline built on diffusion-based text-to-video models, with a MagicAdapter, dynamic… | 59 | 1338 | active |
| IrisRainbowNeko/genshin_auto_fish A Genshin Impact auto-fishing AI built from a YOLOX object detection model (fish and rod landing point localization) and a DQN reinforcemen… | 23 | 4758 | maintenance |
| liwenxi/SWIFT-AI SWIFT-AI is a deep learning system for extremely fast gigapixel-level visual understanding in scientific applications, such as detecting st… | 29 | 1334 | active |
| mapillary/inplace_abn A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing… | 65 | 1333 | stable |
| jonathan-laurent/AlphaZero.jl A generic, simple, and fast Julia implementation of DeepMind's AlphaZero algorithm for training game-playing agents via self-play and MCTS.… | 64 | 1333 | active |
| wenqsun/DimensionX DimensionX is a research framework that generates photorealistic 3D and 4D scenes from a single image using controllable video diffusion mo… | 43 | 1333 | active |
| bytedance/Lance Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained … | 55 | 1329 | active |
| facebookincubator/AITemplate AITemplate is a Python framework that compiles deep neural networks into high-performance CUDA (NVIDIA) or HIP (AMD) C++ code for fast fp16… | 66 | 4724 | maintenance |
| LLaVA-VL/LLaVA-NeXT LLaVA-NeXT is a collection of open large multimodal models (LLaVA-NeXT, LLaVA-Video, LLaVA-OneVision, LLaVA-Critic-R1) that combine vision … | 64 | 4716 | maintenance |
| ACEsuit/mace MACE is a Python library implementing fast and accurate machine learning interatomic potentials using higher-order equivariant message pass… | 89 | 1324 | active |
| agemagician/ProtTrans ProtTrans provides state-of-the-art pre-trained Transformer language models for protein sequences, trained on thousands of GPUs and hundred… | 33 | 1324 | active |
| sjtuytc/UnboundedNeRFPytorch A PyTorch implementation benchmarking state-of-the-art unbounded (large-scale) neural radiance field methods like NeRF++, DVGO, and Block-N… | 23 | 1324 | active |
| tensorflow/lucid Lucid is a collection of infrastructure and tools for research in neural network interpretability, built on TensorFlow 1.x. It provides fea… | 10 | 4704 | maintenance |
| galilai-group/lejepa LeJEPA is a Python framework for scalable, theoretically grounded self-supervised representation learning based on Joint-Embedding Predicti… | 45 | 1322 | active |
| ImprintLab/Medical-SAM-Adapter Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei… | 39 | 1322 | active |
| stared/livelossplot A Python library that draws live training loss and metric plots inside Jupyter Notebooks for Keras, PyTorch, and other deep learning framew… | 91 | 1319 | stable |
| open-edge-platform/geti Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru… | 98 | 1317 | active |
| sicara/easy-few-shot-learning A Python library (easyfsl) with ready-to-use code and tutorial notebooks for few-shot image classification and meta-learning, built on PyTo… | 23 | 1313 | stable |
| yeyupiaoling/VoiceprintRecognition-Pytorch A PyTorch-based voiceprint recognition (speaker recognition) framework implementing models such as ECAPA-TDNN, ResNetSE, ERes2Net, and CAM+… | 58 | 1312 | active |
| ali-vilab/MimicBrush MimicBrush is the official implementation of a zero-shot image editing method that lets users mask a region in a source image and provide a… | 24 | 1311 | active |
| soda-inria/tabicl TabICLv2 is an open-source tabular foundation model that performs classification and regression via a single forward pass through a pre-tra… | 81 | 1309 | active |
| albermax/innvestigate iNNvestigate is a Python toolbox providing a common interface and out-of-the-box implementations of many neural network explanation methods… | 30 | 1309 | active |
| huawei-noah/Efficient-Computing A collection of efficient deep learning methods from Huawei Noah's Ark Lab, covering model compression, knowledge distillation, pruning, qu… | 32 | 1307 | active |
| DAMO-NLP-SG/VideoLLaMA2 VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro… | 25 | 1307 | active |
| luosiallen/latent-consistency-model Official implementation of Latent Consistency Models (LCM), a diffusion-based approach for synthesizing high-resolution images with few-ste… | 27 | 4615 | maintenance |
| STVIR/pysot PySOT is a Python research platform by SenseTime for single object visual tracking, implementing algorithms such as SiamRPN, SiamRPN++, DaS… | 45 | 4600 | maintenance |
| donydchen/mvsplat MVSplat is a PyTorch implementation of an ECCV 2024 Oral model that predicts 3D Gaussians from sparse multi-view images in a single feed-fo… | 61 | 1296 | active |
| robodhruv/visualnav-transformer Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie… | 19 | 1294 | active |
| fundamentalvision/BEVFormer BEVFormer is the official PyTorch implementation of an ECCV 2022 paper that learns bird's-eye-view (BEV) representations from multi-camera … | 23 | 4579 | maintenance |
| braindecode/braindecode Braindecode is an open-source Python toolbox built on PyTorch for decoding raw electrophysiological brain signals such as EEG, ECoG, and ME… | 98 | 1293 | active |
| W2GenAI-Lab/LucidFlux LucidFlux is a caption-free photo-realistic image restoration model built on a large-scale diffusion transformer, released with inference a… | 55 | 1293 | active |
| plaidml/plaidml PlaidML is a portable tensor compiler that enables deep learning on hardware (especially GPUs and embedded devices) not well supported by m… | 10 | 4566 | maintenance |
| RoyalVane/CLAN Official PyTorch implementation of CLAN, a CVPR 2019 (oral) / TPAMI 2022 method for unsupervised domain adaptation in semantic segmentation… | 32 | 1289 | stable |
| TencentQQGYLab/ELLA ELLA is an Efficient Large Language Model Adapter that equips text-to-image diffusion models with LLM-based text understanding via a Timest… | 25 | 1289 | active |
| buoyancy99/diffusion-forcing Official research code for the NeurIPS paper 'Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion', implementing a metho… | 65 | 1288 | active |
| locuslab/TCN PyTorch implementation of Temporal Convolutional Networks (TCN) with benchmarks from the paper 'An Empirical Evaluation of Generic Convolut… | 32 | 4549 | maintenance |
| huanngzh/MV-Adapter MV-Adapter is a plug-and-play adapter that turns pre-trained text-to-image diffusion models (e.g., SDXL, SD2.1) into multi-view consistent … | 34 | 1285 | active |
| Phantom-video/HuMo HuMo is a research model and Python codebase from Tsinghua University and ByteDance for human-centric video generation using collaborative … | 46 | 1283 | active |
| e3nn/e3nn e3nn is a Python/PyTorch library for building E(3)-equivariant neural networks that respect 3D rotation, translation, and mirror symmetries… | 70 | 1280 | active |
| ZhengyiLuo/PHC Official implementation of the ICCV 2023 paper 'Perpetual Humanoid Control for Real-time Simulated Avatars'. It provides a Python codebase … | 44 | 1280 | active |
| AaronJackson/vrn Research code for the ICCV 2017 paper 'Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression'. It uses… | 32 | 4517 | maintenance |
| plemeri/transparent-background A Python tool and CLI that removes backgrounds from images and videos using the InSPyReNet deep learning model (ACCV 2022). It supports ima… | 63 | 1278 | active |
| suragnair/alpha-zero-general A clean, flexible implementation of the AlphaZero self-play reinforcement learning algorithm that can be adapted to any two-player turn-bas… | 32 | 4506 | maintenance |
| studio-dots-ai/dots.tts dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive text-to-speech system, distributed as a Python library with pretrain… | 79 | 1275 | active |
| Ma-Lab-Berkeley/CRATE CRATE is the official PyTorch implementation of the Coding RAte reduction TransformEr, a family of 'white-box' transformer architectures de… | 29 | 1275 | active |
| google-research/simclr Google Research's official implementation of SimCLR and SimCLRv2, a framework for contrastive learning of visual representations, with 65 p… | 10 | 4502 | maintenance |
| Stable-X/Stable3DGen Stable3DGen is a modular Python framework for generating 3D assets from images, adapted from Microsoft's TRELLIS with NVIDIA library depend… | 33 | 1274 | active |
| DreamTechAI/Direct3D-S2 Direct3D-S2 is a research framework for high-resolution 3D shape generation from images, built on sparse volumetric representations and a n… | 29 | 1274 | active |
| dcharatan/pixelsplat pixelSplat is a PyTorch implementation of a feed-forward model that reconstructs 3D radiance fields parameterized by 3D Gaussian primitives… | 27 | 1274 | stable |
| nianticlabs/monodepth2 Monodepth2 is the reference PyTorch implementation of the ICCV 2019 paper 'Digging into Self-Supervised Monocular Depth Prediction'. It tra… | 32 | 4497 | maintenance |
| microsoft/BioGPT BioGPT is Microsoft's domain-specific generative Transformer language model pre-trained on biomedical text, with implementation code and pr… | 32 | 4488 | maintenance |
| NVlabs/stylegan2-ada-pytorch Official PyTorch implementation of StyleGAN2-ADA, a generative adversarial network with adaptive discriminator augmentation for training wi… | 32 | 4487 | maintenance |
| leoxiaobin/deep-high-resolution-net.pytorch Official PyTorch implementation of HRNet (Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019), which maintai… | 32 | 4480 | maintenance |
| luo3300612/Visualizer A lightweight Python library that extracts attention maps and other local variables from deep inside PyTorch models for visualization. It w… | 32 | 1269 | stable |
| lucidrains/flamingo-pytorch A PyTorch implementation of DeepMind's Flamingo visual language model architecture, providing the Perceiver Resampler and Gated Cross-Atten… | 23 | 1269 | active |
| amaiya/ktrain ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul… | 25 | 1268 | active |
| thunlp/OpenNRE OpenNRE is an open-source Python toolkit for neural relation extraction, extracting relation triples between entities from plain text. It u… | 32 | 4467 | maintenance |
| nv-tlabs/Difix3D Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau… | 32 | 1266 | active |
| bryandlee/animegan2-pytorch A PyTorch implementation of AnimeGANv2, a GAN-based image-to-image style transfer model that converts photos into anime-style images. It pr… | 32 | 4452 | maintenance |
| rlawjdghek/StableVITON StableVITON is the official PyTorch implementation of a CVPR 2024 paper that performs image-based virtual try-on using a pre-trained latent… | 47 | 1261 | stable |
| Fictionarry/ER-NeRF ER-NeRF is the official PyTorch implementation of an ICCV 2023 paper on region-aware Neural Radiance Fields for high-fidelity talking portr… | 24 | 1260 | stable |
| nv-tlabs/GET3D GET3D is NVIDIA's PyTorch implementation of a generative model that synthesizes high-quality 3D textured meshes (cars, chairs, animals, bui… | 32 | 4435 | maintenance |
| Francis-Rings/StableAvatar StableAvatar is an end-to-end video diffusion transformer that generates infinite-length, high-quality talking avatar videos from a referen… | 46 | 1258 | active |
| Yuliang-Liu/MonkeyOCRv2 MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2… | 58 | 1256 | active |
| stardist/stardist StarDist is a Python library for object detection and instance segmentation in 2D and 3D microscopy images using star-convex shapes, built … | 60 | 1255 | stable |
| ingra14m/Deformable-3D-Gaussians Official PyTorch implementation of the CVPR 2024 paper 'Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction'. … | 18 | 1255 | stable |
| thu-ml/Motus Motus is the official implementation of a unified latent action world model for robotics, combining a video generation model, a vision-lang… | 43 | 1246 | active |
| LTH14/fractalgen A PyTorch implementation of Fractal Generative Models (FractalGen), enabling pixel-by-pixel high-resolution image generation. It includes p… | 24 | 1244 | active |
| 3DTopia/OpenLRM OpenLRM is an open-source PyTorch implementation of Large Reconstruction Models (LRM) that reconstruct 3D objects (meshes and rendered vide… | 17 | 1244 | active |
| HJYao00/Mulberry Mulberry is a research implementation of an o1-like multimodal large language model (MLLM) that performs step-by-step reasoning and reflect… | 49 | 1243 | active |
| cvg/depthsplat DepthSplat is a PyTorch research library implementing a CVPR 2025 model that connects Gaussian splatting with single/multi-view depth estim… | 56 | 1242 | active |
| showlab/Tune-A-Video Tune-A-Video is the official PyTorch implementation of an ICCV 2023 paper that fine-tunes pre-trained text-to-image diffusion models (like … | 31 | 4364 | maintenance |