Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: gpu-computing

395 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
mujocolab/mjlab
mjlab is a Python framework that combines Isaac Lab's manager-based API with MuJoCo Warp, a GPU-accelerated version of MuJoCo, for reinforc…
852837active
leptonai/leptonai
LeptonAI is a Python library and `lep` CLI for operating NVIDIA DGX Cloud Lepton, a platform that unifies global GPU compute for AI develop…
942825active
huggingface/nanotron
Nanotron is a minimalistic Python library from Hugging Face for pretraining large language models with 3D parallelism (data, tensor, and pi…
562800active
black-forest-labs/flux2
Official inference repository for Black Forest Labs' FLUX.2 family of open-weight image generation and editing models. It provides minimal …
482642active
EMI-Group/evox
EvoX is a distributed GPU-accelerated framework for evolutionary computation, compatible with PyTorch and built on JAX. It provides 50+ evo…
802526active
google-coral/coralnpu
Coral NPU is an open-source neural processing unit (NPU) hardware IP core from Google Research, built on the 32-bit RISC-V ISA with matrix,…
732526active
learning-at-home/hivemind
Hivemind is a PyTorch library for decentralized deep learning across the Internet, enabling training of large models on hundreds of volunte…
592515active
data-infra/cube-studio
CubeStudio is an open-source, cloud-native, all-in-one AI platform covering the full machine learning lifecycle (MLOps/MaaS/LLMOps), includ…
802448active
antirez/h3.c
A native C inference engine for the MiniMax H3 model on Apple Silicon, using Metal for GPU acceleration. It generates video (and audio) fro…
562443active
Alibaba-Quark/LiveAvatar
LiveAvatar is an open-source implementation of an ECCV 2026 paper for streaming, real-time, infinite-length audio-driven avatar video gener…
612386active
microsoft/Olive
Olive is Microsoft's AI model optimization toolkit for the ONNX Runtime, automating finetuning, conversion, quantization, and compression o…
912382active
apple/axlearn
AXLearn is a Python deep learning library built on JAX and XLA for developing and training large-scale models, with an object-oriented conf…
712372active
mflux-community/mflux
MFLUX is a native MLX implementation of state-of-the-art generative image and video models (Flux, Qwen-Image, Z-Image, and others), ported …
902291active
radixark/miles
Miles is an open-source, enterprise-grade reinforcement learning framework for large-scale LLM and VLM post-training, forked from and co-ev…
722263active
rapidsai/cugraph
cuGraph is NVIDIA's RAPIDS collection of GPU-accelerated graph analytics libraries, offering Python, C, and C++ APIs for building graphs an…
942225active
dstackai/dstack
dstack is an open-source, vendor-agnostic control plane for GPU provisioning and orchestration that works across GPU clouds, Kubernetes, an…
952221active
NovaSky-AI/SkyRL
SkyRL is a modular full-stack reinforcement learning library for post-training large language models, combining a training framework (skyrl…
802201active
ByteDance-Seed/VeOmni
VeOmni is a PyTorch-native framework for single- and multi-modal model pre-training and post-training, with a modular, trainer-free design …
822173active
google-deepmind/mujoco_playground
MuJoCo Playground is an open-source Python library of GPU-accelerated robot learning environments built on MuJoCo MJX and MuJoCo Warp. It s…
762168active
fastmachinelearning/hls4ml
hls4ml is a Python package that converts machine learning models from Keras, PyTorch, and ONNX into high-level synthesis (C++) code for FPG…
822114active
vitoplantamura/OnnxStream
A lightweight C++ inference library for ONNX models that streams weights to run large models in very little memory, accelerated by XNNPACK.…
592086active
selkies-project/selkies
Selkies is an open-source low-latency, GPU/CPU-accelerated Linux remote desktop and game streaming platform that delivers an HTML5 web clie…
672067active
marcoslucianops/DeepStream-Yolo
A collection of configuration files, parsers, and conversion utilities for running YOLO-family object detection models on NVIDIA DeepStream…
612054active
deepmodeling/deepmd-kit
DeePMD-kit is a deep learning package for building many-body potential energy representations and running molecular dynamics simulations. I…
952021stable
PrimeIntellect-ai/prime-rl
prime-rl is a Python framework for large-scale, fully asynchronous reinforcement learning training of language models, built on FSDP2 for t…
871975active
NVIDIA-NeMo/RL
NeMo RL is NVIDIA's open-source post-training library for scaling reinforcement learning methods (GRPO, PPO, DPO, SFT, distillation) on LLM…
811961active
openmm/openmm
OpenMM is a high-performance toolkit and library for molecular dynamics simulation, with optimized GPU-accelerated kernels. It can be used …
951960stable
sapientinc/HRM-Text
HRM-Text is a 1B-parameter text generation model based on the hierarchical recurrent HRM architecture, released with a complete pretraining…
531899active
llvm/torch-mlir
Torch-MLIR is a compiler project providing first-class translation of PyTorch programs into the MLIR compiler ecosystem. It lets hardware v…
671892active
laugh12321/TensorRT-YOLO
A C++/Python deployment toolkit for running YOLO-family models (YOLOv3 through YOLO26) on NVIDIA GPUs using TensorRT, with custom plugins, …
631880active
NVIDIA-AI-IOT/Lidar_AI_Solution
NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil…
721867active
nvidia-isaac/cuVSLAM
cuVSLAM is NVIDIA's CUDA-accelerated library for real-time visual odometry and simultaneous localization and mapping (SLAM). It supports mu…
821773active
beam-cloud/beta9
Beam (beta9) is an open-source serverless runtime for AI workloads, providing GPU inference endpoints, isolated sandboxes for running untru…
891755active
kingoflolz/mesh-transformer-jax
A JAX/Haiku library implementing model-parallel training and inference of transformer models using xmap/pjit operators, similar to Megatron…
326380maintenance
gomlx/gomlx
GoMLX is an accelerated machine learning and math framework for Go, comparable to PyTorch/JAX/TensorFlow. It offers differentiable operator…
951621active
UbiquitousLearning/mllm
MLLM is a fast, lightweight multimodal LLM inference engine written in C++ for mobile and edge devices, with backends for ARM CPU, Qualcomm…
731593active
alibaba/Pai-Megatron-Patch
Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and V…
561591active
pytorch/FBGEMM
FBGEMM is a collection of highly optimized low-precision matrix multiplication and convolution kernels for server-side deep learning infere…
931584active
xLLM-AI/xllm
xLLM is a high-performance C++ inference engine for LLM, VLM, DiT and recommendation models, optimized for heterogeneous AI accelerators su…
781536active
mit-han-lab/torchsparse
TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a…
261472active
Lightning-AI/lightning-thunder
Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a …
721469active
FeiYull/TensorRT-Alpha
A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl…
321460active
deepseek-ai/EPLB
EPLB is DeepSeek's open-source Expert Parallelism Load Balancer for Mixture-of-Experts models. It computes balanced expert replication and …
261424active
heterodb/pg-strom
PG-Strom is a PostgreSQL extension that accelerates SQL analytics and batch workloads using GPU devices, NVMe-SSD storage, and Apache Arrow…
761408active
mratsim/Arraymancer
Arraymancer is a fast, ergonomic N-dimensional tensor (ndarray) library written in Nim, inspired by NumPy and PyTorch. It provides CPU, CUD…
611407active
erwincoumans/tiny-differentiable-simulator
Tiny Differentiable Simulator (TDS) is a header-only C++ and CUDA physics library for rigid-body dynamics with zero dependencies, supportin…
231371active
BICLab/SpikingBrain-7B
SpikingBrain-7B is a brain-inspired large language model that combines hybrid efficient attention, MoE modules, and spike encoding, with a …
541369active
uxlfoundation/scikit-learn-intelex
Intel's Extension for Scikit-learn is a free AI accelerator that speeds up existing scikit-learn workflows on CPUs and GPUs, claiming up to…
911356active
KhronosGroup/SPIRV-Tools
SPIR-V Tools is a Khronos Group project providing an API and command-line tools for processing SPIR-V modules, including an assembler, bina…
941353stable
hao-ai-lab/LookaheadDecoding
A Python library implementing Lookahead Decoding, an exact parallel decoding algorithm that accelerates LLM inference without a draft model…
311342active
mapillary/inplace_abn
A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing…
651333stable
plaidml/plaidml
PlaidML is a portable tensor compiler that enables deep learning on hardware (especially GPUs and embedded devices) not well supported by m…
104566maintenance
ModelCloud/GPTQModel
GPTQModel is a production-ready Python toolkit for quantizing (compressing) large language models using GPTQ, AWQ, and related methods, wit…
911248active
ai-dynamo/nixl
NVIDIA Inference Xfer Library (NIXL) is a C++/Python library that accelerates point-to-point communication in AI inference frameworks like …
881227active
Zefan-Cai/R-KV
R-KV is a training-free, redundancy-aware KV cache compression method for reasoning LLMs, discarding repetitive tokens on-the-fly during de…
601209active
dendenxu/fast-gaussian-rasterization
A drop-in replacement for diff-gaussian-rasterization that renders 3D Gaussian Splatting scenes using a geometry-shader-based GPU pipeline …
191202active
Cerebras/modelzoo
Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and …
771193active
higgsfield-ai/higgsfield
Higgsfield is an open-source GPU orchestration and machine learning framework for fault-tolerant, distributed training of very large models…
234106maintenance
steelbrain/ffmpeg-over-ip
A client/server tool that lets applications use GPU-accelerated ffmpeg on a remote machine over a single TCP connection, without GPU passth…
831176active
googlecolab/google-colab-cli
A Python-based command-line interface for Google Colab that lets users provision CPU, GPU, and TPU runtimes, execute local scripts and note…
591126active
MoonshotAI/MoonEP
MoonEP is an Expert Parallelism communication library for Mixture-of-Experts training that keeps token loads perfectly balanced across rank…
561101active
NVIDIA-Omniverse/kit-app-template
A template toolkit from NVIDIA for building GPU-accelerated, OpenUSD-based 3D applications with the Omniverse Kit SDK. It provides pre-conf…
751090active
flagos-ai/FlagGems
FlagGems is a high-performance operator library for large language models written in the Triton language, providing backend-neutral GPU ker…
891089active
NVlabs/Fast-dLLM
NVIDIA's official implementation of Fast-dLLM, a family of training-free and fine-tuning-based acceleration techniques for diffusion-based …
571082active
meta-pytorch/monarch
Monarch is a distributed programming framework for PyTorch built on scalable actor messaging, with actors grouped into meshes, supervision-…
801073active
open-gigaai/giga-train
GigaTrain is an efficient and scalable Python training framework for large AI models, supporting distributed multi-GPU/multi-node execution…
621072active
sirius-db/sirius
Sirius is a GPU-native SQL analytics engine written in C++ that accelerates query execution by offloading it to GPUs. It integrates with ex…
691059active
XYZ-AI-Lab/axrl
AxisRL is an agentic reinforcement learning post-training framework for large language models, built on SGLang for high-throughput rollout …
551056active
Xilinx/finn
FINN is an open-source dataflow compiler from AMD/Xilinx that generates highly efficient FPGA accelerators for quantized neural network (QN…
681046active
aiptimizer/TurboOCR
TurboOCR is an extremely fast GPU-accelerated document parser written in C++ that combines OCR, layout analysis, table extraction, and form…
821043active
scrya-com/rotorquant
RotorQuant is a KV cache compression method for LLM inference that replaces full d×d orthogonal rotations with block-diagonal Clifford roto…
501043active
HannesStark/boltzgen
BoltzGen is an open-source all-atom generative diffusion model for designing protein and peptide binders against arbitrary biomolecular tar…
711042active
NVIDIA-NeMo/Skills
Nemo-Skills is a collection of Python pipelines for improving the skills of large language models, covering synthetic data generation, mode…
701031active
fla-org/native-sparse-attention
Efficient Triton kernel implementations of Native Sparse Attention (NSA), a hardware-aligned, natively trainable sparse attention mechanism…
491020active
microsoft/Tutel
Tutel is Microsoft's optimized Mixture-of-Experts (MoE) library for efficient training and inference of large language models, featuring dy…
791016active
facebookresearch/fairscale
FairScale is a PyTorch extension library providing composable modules and APIs for high-performance, large-scale distributed training, incl…
103407maintenance
isaac-sim/IsaacGymEnvs
A collection of example reinforcement learning environments for NVIDIA Isaac Gym, a GPU-accelerated physics simulator. It provides a Gym-st…
102952maintenance
nerdyrodent/VQGAN-CLIP
A Python application for running VQGAN+CLIP text-to-image generation locally on your own GPU, derived from Katherine Crowson's Google Colab…
322647maintenance
IST-DASLab/gptq
Reference implementation of GPTQ, a one-shot post-training weight quantization method for large generative transformer models based on appr…
322360maintenance
mit-han-lab/once-for-all
Once-for-All (OFA) is a PyTorch library implementing the ICLR 2020 Once-for-All network, which trains a single supernet that can be special…
231956maintenance
Maratyszcza/NNPACK
NNPACK is a C99 acceleration package providing high-performance SIMD and multi-core CPU implementations of neural network layers, especiall…
321710maintenance
bes-dev/stable_diffusion.openvino
A Python CLI implementation of Stable Diffusion text-to-image generation optimized for Intel CPUs and GPUs via OpenVINO. It supports text-t…
321533maintenance
BigScience
BigScience is a research project training large transformer language models (BERT, GPT-style) at scale, built on a fork of Megatron-LM inte…
321448maintenance
chengzeyi/stable-fast
Stable Fast is an ultra lightweight performance optimization framework for HuggingFace Diffusers pipelines on NVIDIA GPUs. It achieves SOTA…
241302maintenance
NVIDIA-AI-IOT/trt_pose
trt_pose is a Python library from NVIDIA for real-time human pose estimation accelerated with TensorRT, targeting NVIDIA Jetson and other N…
231065maintenance
microsoft/nnfusion
NNFusion is a flexible and efficient deep neural network (DNN) compiler that generates high-performance executables from model descriptions…
231002maintenance
Rust-GPU/rust-gpu
A Rust compiler backend that emits SPIR-V, making Rust a first-class language for writing GPU shaders targeting Vulkan. It lets developers …
843308experimental
jephersonRD/Maquina-V5
A collection of Google Colab notebooks and scripts that spin up a temporary cloud gaming PC with an NVIDIA Tesla T4 GPU, install Steam, and…
501763experimental
volcengine/veScale
veScale is a PyTorch distributed training library from ByteDance for hyperscale training of large language models and reinforcement learnin…
571036experimental
horovod/horovod
Horovod is a distributed deep learning training framework for TensorFlow, Keras, PyTorch, and Apache MXNet, originally developed at Uber. I…
1014688abandoned
intel/ipex-llm
IPEX-LLM is a PyTorch LLM acceleration library for Intel hardware (iGPU, NPU, Arc/Flex/Max GPUs, and CPU), offering low-bit quantization (F…
108859abandoned
NVIDIA/DIGITS
DIGITS is a web application for training deep learning models on GPUs, supporting frameworks like Caffe, Torch, and TensorFlow with a brows…
104177abandoned
golemfactory/clay
Clay Golem is a Python implementation of a decentralized peer-to-peer marketplace for renting idle CPU and GPU computing power, with Ethere…
102877abandoned
intel/intel-extension-for-pytorch
A Python package extending official PyTorch with Intel-specific optimizations for CPUs and GPUs, including quantization and LLM inference a…
102012abandoned
nod-ai/AMD-SHARK-Studio
AMD-SHARK Studio is a web UI distribution for high-performance machine learning inference built on SHARK and IREE, primarily for running St…
481451abandoned

← prev page 4 / 4