domain: deep-learning
2771 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| huggingface/swift-coreml-diffusers A native SwiftUI application demonstrating how to run Stable Diffusion text-to-image generation on-device using Apple's Core ML Stable Diff… | 54 | 2757 | active |
| CVCUDA/CV-CUDA CV-CUDA is an open-source GPU-accelerated library of computer vision and image processing operators built on CUDA, with C++ and Python APIs… | 93 | 2723 | active |
| magic-research/magic-animate MagicAnimate is the official implementation of a CVPR 2024 diffusion-based human image animation framework that animates a reference image … | 44 | 10898 | maintenance |
| KimMeen/Time-LLM Time-LLM is the official PyTorch implementation of an ICLR 2024 paper that reprograms frozen large language models (Llama, GPT-2, BERT) for… | 47 | 2689 | active |
| princeton-vl/DROID-SLAM DROID-SLAM is a deep learning-based visual SLAM system for monocular, stereo, and RGB-D cameras that estimates camera trajectory and dense … | 41 | 2678 | active |
| aigc3d/LHM LHM is a PyTorch-based large reconstruction model that reconstructs high-fidelity animatable 3D human avatars from a single image in second… | 52 | 2664 | active |
| HiDream-ai/HiDream-I1 HiDream-I1 is an open-source 17B-parameter text-to-image generative foundation model based on a Sparse Diffusion Transformer, with full and… | 34 | 2515 | active |
| wgsxm/PartCrafter PartCrafter is a structured 3D generative model that jointly generates multiple semantically meaningful 3D mesh parts and objects from a si… | 53 | 2472 | active |
| OpenDCAI/DataFlex DataFlex is a data-centric training framework built on LLaMA-Factory that dynamically schedules LLM training data during optimization. It p… | 67 | 2432 | active |
| HybridRobotics/whole_body_tracking BeyondMimic's official motion tracking training code, a humanoid control framework built on Isaac Lab that trains sim-to-real-ready whole-b… | 69 | 2367 | active |
| NVlabs/ProtoMotions ProtoMotions is a GPU-accelerated simulation and reinforcement learning framework for training physically simulated digital humans and huma… | 66 | 2362 | active |
| NVlabs/nvdiffrec nvdiffrec is NVIDIA's official implementation of a CVPR 2022 oral paper that jointly optimizes triangular 3D meshes, PBR materials, and lig… | 76 | 2296 | stable |
| togethercomputer/OpenChatKit OpenChatKit is an open-source toolkit from Together, LAION, and Ontocord.ai for building and fine-tuning chat language models, including in… | 30 | 8983 | maintenance |
| nv-tlabs/lyra Project Lyra is NVIDIA's open series of generative 3D world models, including Lyra 1.0 for feed-forward 3D/4D scene generation from a singl… | 59 | 2279 | active |
| langfengQ/verl-agent verl-agent is an extension of the veRL framework for training LLM and VLM agents via reinforcement learning, featuring step-independent mul… | 56 | 2276 | active |
| Liuziyu77/Visual-RFT Official research code for Visual-RFT and Visual-ARFT, applying GRPO-based reinforcement fine-tuning with rule-based verifiable rewards to … | 42 | 2274 | active |
| OlafenwaMoses/ImageAI ImageAI is a Python library that lets developers add computer vision capabilities like image classification, object detection, and video ob… | 23 | 8882 | maintenance |
| jd-opensource/JoyAI-Image JoyAI-Image is a unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editin… | 58 | 2149 | active |
| PKU-YuanGroup/LLaVA-CoT LLaVA-CoT is an 11B vision language model with training, inference, and dataset-generation code for spontaneous step-by-step multimodal rea… | 47 | 2132 | active |
| fastmachinelearning/hls4ml hls4ml is a Python package that converts machine learning models from Keras, PyTorch, and ONNX into high-level synthesis (C++) code for FPG… | 82 | 2124 | active |
| facebookresearch/theseus Theseus is a PyTorch-based library for building custom differentiable nonlinear optimization layers, supporting problems in robotics and vi… | 23 | 2056 | active |
| serengil/retinaface RetinaFace is a Python library for deep learning based face detection, built on TensorFlow and derived from the insightface project's Retin… | 61 | 2032 | active |
| sicxu/Deep3DFaceRecon_pytorch A PyTorch implementation of Deep3DFaceReconstruction, a weakly-supervised CNN method for reconstructing 3D face geometry from a single imag… | 32 | 1911 | stable |
| diffgram/diffgram Diffgram is a self-hosted AI datastore for managing schemas, BLOBs, and predictions, with built-in human supervision (data labeling), data … | 62 | 1909 | active |
| PRIME-RL/SimpleVLA-RL SimpleVLA-RL is an open-source reinforcement learning framework for training Vision-Language-Action (VLA) models for robotic manipulation, … | 46 | 1845 | active |
| YangLing0818/RPG-DiffusionMaster Official implementation of RPG (Recaption, Plan, Generate), a training-free framework that uses multimodal LLMs as prompt recaptioners and … | 28 | 1844 | active |
| jd-opensource/JoyAI-VL-Interaction JoyAI-VL-Interaction is an open 8B-scale vision-language interaction model with a complete deployable real-time streaming system, including… | 58 | 1840 | active |
| Zheng-Chong/CatVTON CatVTON is a lightweight diffusion model for virtual try-on that swaps clothing onto a person image using a concatenation-based architectur… | 40 | 1830 | active |
| Robbyant/lingbot-vla LingBot-VLA is a Vision-Language-Action foundation model for robot manipulation, pretrained on 20,000 hours of real-world dual-arm robot da… | 54 | 1794 | active |
| Emu Series Emu3 is a suite of state-of-the-art multimodal models from BAAI trained solely with next-token prediction, tokenizing images, text, and vid… | 57 | 1779 | active |
| GAP-LAB-CUHK-SZ/gaustudio GauStudio is a modular PyTorch framework for 3D Gaussian Splatting (3DGS) research and development, supporting novel view synthesis, 3D rec… | 49 | 1762 | active |
| One-2-3-45/One-2-3-45 One-2-3-45 is the official PyTorch implementation of a NeurIPS 2023 paper that converts any single image into a full 360-degree 3D textured… | 29 | 1718 | stable |
| williamyang1991/DualStyleGAN Official PyTorch implementation of DualStyleGAN, a CVPR 2022 model for exemplar-based high-resolution (1024px) portrait style transfer. It … | 32 | 1683 | stable |
| Stability-AI/stable-virtual-camera Stable Virtual Camera (SEVA) is a generalist diffusion model for novel view synthesis that generates 3D-consistent views of a scene from an… | 54 | 1657 | active |
| ZJU4HealthCare/HealthGPT HealthGPT is a medical multimodal large language model family unifying medical image comprehension and generation via heterogeneous knowled… | 63 | 1655 | active |
| guochengqian/Magic123 Magic123 is the official PyTorch implementation of an ICLR 2024 paper that generates high-quality textured 3D meshes from a single unposed … | 90 | 1623 | stable |
| PKU-Alignment/safe-rlhf Beaver is a modular open-source framework from Peking University for training large language models with SFT, RLHF, and Safe RLHF (constrai… | 53 | 1613 | active |
| zju3dv/EasyVolcap EasyVolcap is a PyTorch-based library for accelerating neural volumetric video research, covering volumetric video capture, reconstruction,… | 27 | 1587 | active |
| photosynthesis-team/piq PyTorch Image Quality (PIQ) is a collection of measures and metrics for image quality assessment in image-to-image tasks, written in pure P… | 23 | 1574 | stable |
| tianrun-chen/SAM-Adapter-PyTorch A PyTorch library that adapts Meta AI's Segment Anything Model (SAM, SAM2, SAM3) to underperforming downstream segmentation tasks using lig… | 67 | 1551 | active |
| RightNow-AI/autokernel AutoKernel is an open-source autoresearch pipeline that takes any PyTorch model, profiles it to find GPU kernel bottlenecks, extracts them … | 48 | 1548 | active |
| luciddreamer-cvlab/LucidDreamer LucidDreamer is the official implementation of a research method that generates 3D Gaussian Splatting scenes from text prompts, published i… | 69 | 1528 | active |
| ATH-MaaS/Ovis Ovis is an open-source Multimodal Large Language Model (MLLM) architecture that structurally aligns visual and textual embeddings, with rel… | 65 | 1514 | active |
| hkchengrex/Tracking-Anything-with-DEVA DEVA is a decoupled video segmentation framework that combines task-specific image-level segmentation models with a universal bi-directiona… | 27 | 1509 | stable |
| Vincentqyw/cv-arxiv-daily An automated daily digest of computer vision and robotics arXiv papers (SLAM, SFM, visual localization, keypoint detection, image matching,… | 77 | 1495 | active |
| CSAILVision/semantic-segmentation-pytorch A PyTorch implementation of semantic segmentation (scene parsing) models for the MIT ADE20K dataset, including pretrained model zoo and tra… | 32 | 5080 | maintenance |
| alibaba/Logics-Parsing Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul… | 54 | 1404 | active |
| lxtGH/OMG-Seg Official research codebase for OMG-Seg (CVPR 2024) and OMG-LLaVA (NeurIPS 2024), unified models for image-level, object-level, and pixel-le… | 47 | 1354 | active |
| CarperAI/trlx trlX is a distributed training framework for fine-tuning large language models with reinforcement learning from human feedback (RLHF), supp… | 23 | 4753 | maintenance |
| Henry-23/VideoChat A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-late… | 48 | 1303 | active |
| deepseek-ai/DeepSeek-Prover-V2 DeepSeek-Prover-V2 is an open-source large language model for formal theorem proving in Lean 4, trained via reinforcement learning with sub… | 34 | 1298 | active |
| donydchen/mvsplat MVSplat is a PyTorch implementation of an ECCV 2024 Oral model that predicts 3D Gaussians from sparse multi-view images in a single feed-fo… | 61 | 1296 | active |
| BytedTsinghua-SIA/CUDA-Agent CUDA-Agent is a large-scale agentic reinforcement learning system from ByteDance Seed and Tsinghua that trains LLMs to generate high-perfor… | 56 | 1290 | active |
| zju3dv/MatchAnything MatchAnything is a deep learning model for universal cross-modality image matching, released as research code accompanying a TPAMI 2026 pap… | 64 | 1284 | active |
| Visual-Agent/DeepEyes DeepEyes is a research project that trains multimodal vision-language models to 'think with images' using end-to-end reinforcement learning… | 44 | 1273 | active |
| nv-tlabs/Difix3D Difix3D+ is a research codebase from NVIDIA implementing a single-step diffusion model pipeline that removes artifacts from NeRF and 3D Gau… | 32 | 1268 | active |
| rlawjdghek/StableVITON StableVITON is the official PyTorch implementation of a CVPR 2024 paper that performs image-based virtual try-on using a pre-trained latent… | 47 | 1263 | stable |
| ziyc/drivestudio DriveStudio is a Python framework for 3D Gaussian Splatting (3DGS) based reconstruction and simulation of dynamic urban driving scenes. It … | 40 | 1256 | active |
| VainF/pytorch-msssim A PyTorch library providing fast, differentiable SSIM and MS-SSIM image quality metrics using separable Gaussian filtering for speed. It ca… | 23 | 1253 | stable |
| X-Square-Robot/wall-x Wall-X is the open-source training and inference stack for X Square Robot's WALL series of embodied foundation models (VLAs) for general-pu… | 61 | 1252 | active |
| Linketic/CityGaussian Official implementation of the CityGaussian series (ECCV 2024, ICLR 2025) for high-quality large-scale 3D scene reconstruction with Gaussia… | 66 | 1251 | active |
| unum-cloud/UForm UForm is a compact multimodal AI library providing tiny image-text embedding models (64-768 dimensions, Matryoshka-style) and small generat… | 55 | 1248 | active |
| 3DTopia/OpenLRM OpenLRM is an open-source PyTorch implementation of Large Reconstruction Models (LRM) that reconstruct 3D objects (meshes and rendered vide… | 17 | 1246 | active |
| cvg/depthsplat DepthSplat is a PyTorch research library implementing a CVPR 2025 model that connects Gaussian splatting with single/multi-view depth estim… | 56 | 1243 | active |
| marian42/mesh_to_sdf A Python library that computes approximate signed distance fields (SDFs) for arbitrary triangle meshes, including non-watertight, self-inte… | 32 | 1242 | stable |
| ZHKKKe/MODNet MODNet is a deep learning model for real-time portrait matting (background removal) that requires only an RGB image as input, with no trima… | 32 | 4359 | maintenance |
| apchenstu/TensoRF TensoRF is a PyTorch implementation of the ECCV 2022 paper 'TensoRF: Tensorial Radiance Fields', which models and reconstructs radiance fie… | 44 | 1237 | stable |
| EvolvingLMMs-Lab/LLaVA-OneVision-2 A fully open framework for training multimodal large language models, releasing models, datasets, and training recipes for the LLaVA-OneVis… | 72 | 1197 | active |
| deepseek-ai/DeepSeek-VL DeepSeek-VL is an open-source vision-language foundation model for real-world multimodal understanding, released with model weights and inf… | 25 | 4177 | maintenance |
| Tencent-Hunyuan/HunyuanWorld-Mirror HunyuanWorld-Mirror is a feed-forward 3D reconstruction model from Tencent that predicts camera poses, intrinsics, depth maps, point clouds… | 54 | 1195 | active |
| trianglesplatting/triangle-splatting Official implementation of 'Triangle Splatting for Real-Time Radiance Field Rendering' (3DV 2026), which uses 3D triangles as rendering pri… | 43 | 1193 | active |
| ubicomplab/rPPG-Toolbox rPPG-Toolbox is an open-source Python toolbox for camera-based physiological sensing (remote photoplethysmography), enabling heart rate and… | 51 | 1185 | active |
| chensjtu/GaussianObject GaussianObject is a research framework for high-quality 3D object reconstruction from as few as four input images using Gaussian splatting,… | 26 | 1184 | active |
| DAMO-NLP-SG/VideoLLaMA3 VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and de… | 37 | 1178 | active |
| majianjia/nnom NNoM is a high-level neural network inference library written in C for microcontrollers. It converts Keras models into optimized on-device … | 23 | 1164 | stable |
| TJU-Aerial-Robotics/YOPO YOPO is a learning-based one-stage planner for quadrotor autonomous navigation in obstacle-dense environments, integrating perception, mapp… | 79 | 1159 | active |
| OpenGVLab/VisionLLM VisionLLM is a series of open-source multimodal large language models from OpenGVLab that unify vision-centric tasks under language instruc… | 33 | 1154 | active |
| princeton-vl/DPVO DPVO is a deep learning-based visual odometry and SLAM system that estimates camera trajectories from video or image sequences using patch-… | 32 | 1114 | active |
| jina-ai/discoart DiscoArt is a Python library that wraps Disco Diffusion (CLIP-guided diffusion) so generative artists and developers can create AI artworks… | 23 | 3827 | maintenance |
| szymanowiczs/splatter-image Official PyTorch implementation of 'Splatter Image: Ultra-Fast Single-View 3D Reconstruction' (CVPR 2024), which uses an image-to-image net… | 26 | 1105 | active |
| yzslab/gaussian-splatting-lightning A PyTorch Lightning implementation of 3D Gaussian Splatting with many derived algorithms (Mip-Splatting, LightGaussian, 2DGS, deformable Ga… | 58 | 1103 | active |
| Toni-SM/skrl skrl is an open-source modular Reinforcement Learning library written in Python, implemented in PyTorch, JAX, and NVIDIA Warp. It supports … | 81 | 1095 | active |
| LujiaJin/One-Pot_Multi-Frame_Denoising Official PyTorch implementation of the One-Pot Multi-frame Denoising (OPD) method published at BMVC 2022 and extended in IJCV. It provides … | 60 | 1094 | stable |
| lizhe00/AnimatableGaussians Official PyTorch implementation of the CVPR 2024 paper 'Animatable Gaussians', which learns pose-dependent Gaussian maps for high-fidelity … | 27 | 1094 | active |
| mlc-ai/web-stable-diffusion A project that compiles and runs Stable Diffusion text-to-image models entirely inside web browsers using WebGPU and WebAssembly, with no s… | 30 | 3723 | maintenance |
| Stability-AI/stable-point-aware-3d SPAR3D is Stability AI's open-source model for fast single-image 3D mesh reconstruction using a two-stage pipeline with point cloud conditi… | 30 | 1070 | active |
| X-LANCE/SLAM-LLM SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc… | 55 | 1057 | active |
| ShenhanQian/GaussianAvatars Official research code for GaussianAvatars, a CVPR 2024 Highlight method that creates photorealistic, fully controllable head avatars by ri… | 56 | 1054 | active |
| InternRobotics/PointLLM PointLLM is a multimodal large language model that understands colored 3D point clouds of objects, built on a point cloud encoder fused wit… | 65 | 1053 | active |
| HannesStark/boltzgen BoltzGen is an open-source all-atom generative diffusion model for designing protein and peptide binders against arbitrary biomolecular tar… | 71 | 1049 | active |
| ml-tooling/ml-workspace ML Workspace is an all-in-one web-based IDE Docker image specialized for machine learning and data science. It bundles Jupyter, JupyterLab,… | 23 | 3544 | maintenance |
| rl-tools/rl-tools RLtools is a pure C++ header-only, dependency-free deep reinforcement learning library supporting algorithms like SAC, TD3, and PPO. It com… | 65 | 1033 | active |
| HarborYuan/ovsam Official PyTorch implementation of Open-Vocabulary SAM (ECCV 2024), a model that unifies SAM's interactive segmentation with CLIP's open-vo… | 42 | 1033 | active |
| JackAILab/ConsistentID ConsistentID is a diffusion-based portrait generation model and toolkit that preserves facial identity from a single reference image using … | 52 | 1026 | active |
| yangxy/PASD PASD (Pixel-Aware Stable Diffusion) is a Python research codebase implementing an ECCV 2024 method for realistic image super-resolution and… | 28 | 1022 | active |
| Alpha-VLLM/Lumina-DiMOO Lumina-DiMOO is an open-source omni diffusion large language model that uses fully discrete diffusion to handle multimodal inputs and outpu… | 55 | 1020 | active |
| SunOner/sunone_aimbot An AI-powered aimbot for first-person shooter games that uses YOLO object detection models (YOLOv8/v10/v12) with TensorRT/ONNX acceleration… | 65 | 1014 | active |
| TRI-ML/prismatic-vlms Prismatic VLMs is a PyTorch-based codebase for training visually-conditioned language models (VLMs) with flexible vision backbones like CLI… | 25 | 1013 | active |
| MeiGen-AI/PosterCraft PosterCraft is a unified framework for generating high-quality aesthetic posters, published as an ICLR 2026 paper. It provides model weight… | 48 | 1009 | active |
| TinyLLaVA/TinyLLaVA_Factory TinyLLaVA Factory is an open-source modular PyTorch/HuggingFace codebase for training small-scale large multimodal models (LMMs) that combi… | 68 | 1004 | active |