domain: gpu-computing
395 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| mujocolab/mjlab mjlab is a Python framework that combines Isaac Lab's manager-based API with MuJoCo Warp, a GPU-accelerated version of MuJoCo, for reinforc… | 85 | 2837 | active |
| leptonai/leptonai LeptonAI is a Python library and `lep` CLI for operating NVIDIA DGX Cloud Lepton, a platform that unifies global GPU compute for AI develop… | 94 | 2825 | active |
| huggingface/nanotron Nanotron is a minimalistic Python library from Hugging Face for pretraining large language models with 3D parallelism (data, tensor, and pi… | 56 | 2800 | active |
| black-forest-labs/flux2 Official inference repository for Black Forest Labs' FLUX.2 family of open-weight image generation and editing models. It provides minimal … | 48 | 2642 | active |
| EMI-Group/evox EvoX is a distributed GPU-accelerated framework for evolutionary computation, compatible with PyTorch and built on JAX. It provides 50+ evo… | 80 | 2526 | active |
| google-coral/coralnpu Coral NPU is an open-source neural processing unit (NPU) hardware IP core from Google Research, built on the 32-bit RISC-V ISA with matrix,… | 73 | 2526 | active |
| learning-at-home/hivemind Hivemind is a PyTorch library for decentralized deep learning across the Internet, enabling training of large models on hundreds of volunte… | 59 | 2515 | active |
| data-infra/cube-studio CubeStudio is an open-source, cloud-native, all-in-one AI platform covering the full machine learning lifecycle (MLOps/MaaS/LLMOps), includ… | 80 | 2448 | active |
| antirez/h3.c A native C inference engine for the MiniMax H3 model on Apple Silicon, using Metal for GPU acceleration. It generates video (and audio) fro… | 56 | 2443 | active |
| Alibaba-Quark/LiveAvatar LiveAvatar is an open-source implementation of an ECCV 2026 paper for streaming, real-time, infinite-length audio-driven avatar video gener… | 61 | 2386 | active |
| microsoft/Olive Olive is Microsoft's AI model optimization toolkit for the ONNX Runtime, automating finetuning, conversion, quantization, and compression o… | 91 | 2382 | active |
| apple/axlearn AXLearn is a Python deep learning library built on JAX and XLA for developing and training large-scale models, with an object-oriented conf… | 71 | 2372 | active |
| mflux-community/mflux MFLUX is a native MLX implementation of state-of-the-art generative image and video models (Flux, Qwen-Image, Z-Image, and others), ported … | 90 | 2291 | active |
| radixark/miles Miles is an open-source, enterprise-grade reinforcement learning framework for large-scale LLM and VLM post-training, forked from and co-ev… | 72 | 2263 | active |
| rapidsai/cugraph cuGraph is NVIDIA's RAPIDS collection of GPU-accelerated graph analytics libraries, offering Python, C, and C++ APIs for building graphs an… | 94 | 2225 | active |
| dstackai/dstack dstack is an open-source, vendor-agnostic control plane for GPU provisioning and orchestration that works across GPU clouds, Kubernetes, an… | 95 | 2221 | active |
| NovaSky-AI/SkyRL SkyRL is a modular full-stack reinforcement learning library for post-training large language models, combining a training framework (skyrl… | 80 | 2201 | active |
| ByteDance-Seed/VeOmni VeOmni is a PyTorch-native framework for single- and multi-modal model pre-training and post-training, with a modular, trainer-free design … | 82 | 2173 | active |
| google-deepmind/mujoco_playground MuJoCo Playground is an open-source Python library of GPU-accelerated robot learning environments built on MuJoCo MJX and MuJoCo Warp. It s… | 76 | 2168 | active |
| fastmachinelearning/hls4ml hls4ml is a Python package that converts machine learning models from Keras, PyTorch, and ONNX into high-level synthesis (C++) code for FPG… | 82 | 2114 | active |
| vitoplantamura/OnnxStream A lightweight C++ inference library for ONNX models that streams weights to run large models in very little memory, accelerated by XNNPACK.… | 59 | 2086 | active |
| selkies-project/selkies Selkies is an open-source low-latency, GPU/CPU-accelerated Linux remote desktop and game streaming platform that delivers an HTML5 web clie… | 67 | 2067 | active |
| marcoslucianops/DeepStream-Yolo A collection of configuration files, parsers, and conversion utilities for running YOLO-family object detection models on NVIDIA DeepStream… | 61 | 2054 | active |
| deepmodeling/deepmd-kit DeePMD-kit is a deep learning package for building many-body potential energy representations and running molecular dynamics simulations. I… | 95 | 2021 | stable |
| PrimeIntellect-ai/prime-rl prime-rl is a Python framework for large-scale, fully asynchronous reinforcement learning training of language models, built on FSDP2 for t… | 87 | 1975 | active |
| NVIDIA-NeMo/RL NeMo RL is NVIDIA's open-source post-training library for scaling reinforcement learning methods (GRPO, PPO, DPO, SFT, distillation) on LLM… | 81 | 1961 | active |
| openmm/openmm OpenMM is a high-performance toolkit and library for molecular dynamics simulation, with optimized GPU-accelerated kernels. It can be used … | 95 | 1960 | stable |
| sapientinc/HRM-Text HRM-Text is a 1B-parameter text generation model based on the hierarchical recurrent HRM architecture, released with a complete pretraining… | 53 | 1899 | active |
| llvm/torch-mlir Torch-MLIR is a compiler project providing first-class translation of PyTorch programs into the MLIR compiler ecosystem. It lets hardware v… | 67 | 1892 | active |
| laugh12321/TensorRT-YOLO A C++/Python deployment toolkit for running YOLO-family models (YOLOv3 through YOLO26) on NVIDIA GPUs using TensorRT, with custom plugins, … | 63 | 1880 | active |
| NVIDIA-AI-IOT/Lidar_AI_Solution NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil… | 72 | 1867 | active |
| nvidia-isaac/cuVSLAM cuVSLAM is NVIDIA's CUDA-accelerated library for real-time visual odometry and simultaneous localization and mapping (SLAM). It supports mu… | 82 | 1773 | active |
| beam-cloud/beta9 Beam (beta9) is an open-source serverless runtime for AI workloads, providing GPU inference endpoints, isolated sandboxes for running untru… | 89 | 1755 | active |
| kingoflolz/mesh-transformer-jax A JAX/Haiku library implementing model-parallel training and inference of transformer models using xmap/pjit operators, similar to Megatron… | 32 | 6380 | maintenance |
| gomlx/gomlx GoMLX is an accelerated machine learning and math framework for Go, comparable to PyTorch/JAX/TensorFlow. It offers differentiable operator… | 95 | 1621 | active |
| UbiquitousLearning/mllm MLLM is a fast, lightweight multimodal LLM inference engine written in C++ for mobile and edge devices, with backends for ARM CPU, Qualcomm… | 73 | 1593 | active |
| alibaba/Pai-Megatron-Patch Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and V… | 56 | 1591 | active |
| pytorch/FBGEMM FBGEMM is a collection of highly optimized low-precision matrix multiplication and convolution kernels for server-side deep learning infere… | 93 | 1584 | active |
| xLLM-AI/xllm xLLM is a high-performance C++ inference engine for LLM, VLM, DiT and recommendation models, optimized for heterogeneous AI accelerators su… | 78 | 1536 | active |
| mit-han-lab/torchsparse TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a… | 26 | 1472 | active |
| Lightning-AI/lightning-thunder Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a … | 72 | 1469 | active |
| FeiYull/TensorRT-Alpha A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl… | 32 | 1460 | active |
| deepseek-ai/EPLB EPLB is DeepSeek's open-source Expert Parallelism Load Balancer for Mixture-of-Experts models. It computes balanced expert replication and … | 26 | 1424 | active |
| heterodb/pg-strom PG-Strom is a PostgreSQL extension that accelerates SQL analytics and batch workloads using GPU devices, NVMe-SSD storage, and Apache Arrow… | 76 | 1408 | active |
| mratsim/Arraymancer Arraymancer is a fast, ergonomic N-dimensional tensor (ndarray) library written in Nim, inspired by NumPy and PyTorch. It provides CPU, CUD… | 61 | 1407 | active |
| erwincoumans/tiny-differentiable-simulator Tiny Differentiable Simulator (TDS) is a header-only C++ and CUDA physics library for rigid-body dynamics with zero dependencies, supportin… | 23 | 1371 | active |
| BICLab/SpikingBrain-7B SpikingBrain-7B is a brain-inspired large language model that combines hybrid efficient attention, MoE modules, and spike encoding, with a … | 54 | 1369 | active |
| uxlfoundation/scikit-learn-intelex Intel's Extension for Scikit-learn is a free AI accelerator that speeds up existing scikit-learn workflows on CPUs and GPUs, claiming up to… | 91 | 1356 | active |
| KhronosGroup/SPIRV-Tools SPIR-V Tools is a Khronos Group project providing an API and command-line tools for processing SPIR-V modules, including an assembler, bina… | 94 | 1353 | stable |
| hao-ai-lab/LookaheadDecoding A Python library implementing Lookahead Decoding, an exact parallel decoding algorithm that accelerates LLM inference without a draft model… | 31 | 1342 | active |
| mapillary/inplace_abn A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing… | 65 | 1333 | stable |
| plaidml/plaidml PlaidML is a portable tensor compiler that enables deep learning on hardware (especially GPUs and embedded devices) not well supported by m… | 10 | 4566 | maintenance |
| ModelCloud/GPTQModel GPTQModel is a production-ready Python toolkit for quantizing (compressing) large language models using GPTQ, AWQ, and related methods, wit… | 91 | 1248 | active |
| ai-dynamo/nixl NVIDIA Inference Xfer Library (NIXL) is a C++/Python library that accelerates point-to-point communication in AI inference frameworks like … | 88 | 1227 | active |
| Zefan-Cai/R-KV R-KV is a training-free, redundancy-aware KV cache compression method for reasoning LLMs, discarding repetitive tokens on-the-fly during de… | 60 | 1209 | active |
| dendenxu/fast-gaussian-rasterization A drop-in replacement for diff-gaussian-rasterization that renders 3D Gaussian Splatting scenes using a geometry-shader-based GPU pipeline … | 19 | 1202 | active |
| Cerebras/modelzoo Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and … | 77 | 1193 | active |
| higgsfield-ai/higgsfield Higgsfield is an open-source GPU orchestration and machine learning framework for fault-tolerant, distributed training of very large models… | 23 | 4106 | maintenance |
| steelbrain/ffmpeg-over-ip A client/server tool that lets applications use GPU-accelerated ffmpeg on a remote machine over a single TCP connection, without GPU passth… | 83 | 1176 | active |
| googlecolab/google-colab-cli A Python-based command-line interface for Google Colab that lets users provision CPU, GPU, and TPU runtimes, execute local scripts and note… | 59 | 1126 | active |
| MoonshotAI/MoonEP MoonEP is an Expert Parallelism communication library for Mixture-of-Experts training that keeps token loads perfectly balanced across rank… | 56 | 1101 | active |
| NVIDIA-Omniverse/kit-app-template A template toolkit from NVIDIA for building GPU-accelerated, OpenUSD-based 3D applications with the Omniverse Kit SDK. It provides pre-conf… | 75 | 1090 | active |
| flagos-ai/FlagGems FlagGems is a high-performance operator library for large language models written in the Triton language, providing backend-neutral GPU ker… | 89 | 1089 | active |
| NVlabs/Fast-dLLM NVIDIA's official implementation of Fast-dLLM, a family of training-free and fine-tuning-based acceleration techniques for diffusion-based … | 57 | 1082 | active |
| meta-pytorch/monarch Monarch is a distributed programming framework for PyTorch built on scalable actor messaging, with actors grouped into meshes, supervision-… | 80 | 1073 | active |
| open-gigaai/giga-train GigaTrain is an efficient and scalable Python training framework for large AI models, supporting distributed multi-GPU/multi-node execution… | 62 | 1072 | active |
| sirius-db/sirius Sirius is a GPU-native SQL analytics engine written in C++ that accelerates query execution by offloading it to GPUs. It integrates with ex… | 69 | 1059 | active |
| XYZ-AI-Lab/axrl AxisRL is an agentic reinforcement learning post-training framework for large language models, built on SGLang for high-throughput rollout … | 55 | 1056 | active |
| Xilinx/finn FINN is an open-source dataflow compiler from AMD/Xilinx that generates highly efficient FPGA accelerators for quantized neural network (QN… | 68 | 1046 | active |
| aiptimizer/TurboOCR TurboOCR is an extremely fast GPU-accelerated document parser written in C++ that combines OCR, layout analysis, table extraction, and form… | 82 | 1043 | active |
| scrya-com/rotorquant RotorQuant is a KV cache compression method for LLM inference that replaces full d×d orthogonal rotations with block-diagonal Clifford roto… | 50 | 1043 | active |
| HannesStark/boltzgen BoltzGen is an open-source all-atom generative diffusion model for designing protein and peptide binders against arbitrary biomolecular tar… | 71 | 1042 | active |
| NVIDIA-NeMo/Skills Nemo-Skills is a collection of Python pipelines for improving the skills of large language models, covering synthetic data generation, mode… | 70 | 1031 | active |
| fla-org/native-sparse-attention Efficient Triton kernel implementations of Native Sparse Attention (NSA), a hardware-aligned, natively trainable sparse attention mechanism… | 49 | 1020 | active |
| microsoft/Tutel Tutel is Microsoft's optimized Mixture-of-Experts (MoE) library for efficient training and inference of large language models, featuring dy… | 79 | 1016 | active |
| facebookresearch/fairscale FairScale is a PyTorch extension library providing composable modules and APIs for high-performance, large-scale distributed training, incl… | 10 | 3407 | maintenance |
| isaac-sim/IsaacGymEnvs A collection of example reinforcement learning environments for NVIDIA Isaac Gym, a GPU-accelerated physics simulator. It provides a Gym-st… | 10 | 2952 | maintenance |
| nerdyrodent/VQGAN-CLIP A Python application for running VQGAN+CLIP text-to-image generation locally on your own GPU, derived from Katherine Crowson's Google Colab… | 32 | 2647 | maintenance |
| IST-DASLab/gptq Reference implementation of GPTQ, a one-shot post-training weight quantization method for large generative transformer models based on appr… | 32 | 2360 | maintenance |
| mit-han-lab/once-for-all Once-for-All (OFA) is a PyTorch library implementing the ICLR 2020 Once-for-All network, which trains a single supernet that can be special… | 23 | 1956 | maintenance |
| Maratyszcza/NNPACK NNPACK is a C99 acceleration package providing high-performance SIMD and multi-core CPU implementations of neural network layers, especiall… | 32 | 1710 | maintenance |
| bes-dev/stable_diffusion.openvino A Python CLI implementation of Stable Diffusion text-to-image generation optimized for Intel CPUs and GPUs via OpenVINO. It supports text-t… | 32 | 1533 | maintenance |
| BigScience BigScience is a research project training large transformer language models (BERT, GPT-style) at scale, built on a fork of Megatron-LM inte… | 32 | 1448 | maintenance |
| chengzeyi/stable-fast Stable Fast is an ultra lightweight performance optimization framework for HuggingFace Diffusers pipelines on NVIDIA GPUs. It achieves SOTA… | 24 | 1302 | maintenance |
| NVIDIA-AI-IOT/trt_pose trt_pose is a Python library from NVIDIA for real-time human pose estimation accelerated with TensorRT, targeting NVIDIA Jetson and other N… | 23 | 1065 | maintenance |
| microsoft/nnfusion NNFusion is a flexible and efficient deep neural network (DNN) compiler that generates high-performance executables from model descriptions… | 23 | 1002 | maintenance |
| Rust-GPU/rust-gpu A Rust compiler backend that emits SPIR-V, making Rust a first-class language for writing GPU shaders targeting Vulkan. It lets developers … | 84 | 3308 | experimental |
| jephersonRD/Maquina-V5 A collection of Google Colab notebooks and scripts that spin up a temporary cloud gaming PC with an NVIDIA Tesla T4 GPU, install Steam, and… | 50 | 1763 | experimental |
| volcengine/veScale veScale is a PyTorch distributed training library from ByteDance for hyperscale training of large language models and reinforcement learnin… | 57 | 1036 | experimental |
| horovod/horovod Horovod is a distributed deep learning training framework for TensorFlow, Keras, PyTorch, and Apache MXNet, originally developed at Uber. I… | 10 | 14688 | abandoned |
| intel/ipex-llm IPEX-LLM is a PyTorch LLM acceleration library for Intel hardware (iGPU, NPU, Arc/Flex/Max GPUs, and CPU), offering low-bit quantization (F… | 10 | 8859 | abandoned |
| NVIDIA/DIGITS DIGITS is a web application for training deep learning models on GPUs, supporting frameworks like Caffe, Torch, and TensorFlow with a brows… | 10 | 4177 | abandoned |
| golemfactory/clay Clay Golem is a Python implementation of a decentralized peer-to-peer marketplace for renting idle CPU and GPU computing power, with Ethere… | 10 | 2877 | abandoned |
| intel/intel-extension-for-pytorch A Python package extending official PyTorch with Intel-specific optimizations for CPUs and GPUs, including quantization and LLM inference a… | 10 | 2012 | abandoned |
| nod-ai/AMD-SHARK-Studio AMD-SHARK Studio is a web UI distribution for high-performance machine learning inference built on SHARK and IREE, primarily for running St… | 48 | 1451 | abandoned |
← prev page 4 / 4