Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: gpu-computing

559 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
wilicc/gpu-burn
GPU Burn is a multi-GPU CUDA stress test tool that pushes NVIDIA GPUs to maximum load for stability and thermal testing. It supports config…
702322active
apache/mahout
Apache Mahout is an Apache Software Foundation project providing Qumat, a high-level Python library for quantum computing with a unified AP…
862304active
CoinCheung/pytorch-loss
A PyTorch library providing a collection of loss functions (focal loss, triplet loss, AMSoftmax, label-smooth CE, dice loss, lovasz-softmax…
322252active
kubeflow/trainer
Kubeflow Trainer is a Kubernetes-native platform for distributed AI model training and LLM fine-tuning across frameworks like PyTorch, JAX,…
952198active
NVIDIA/cutile-python
cuTile Python is a tile-based programming language and library from NVIDIA for writing parallel kernels that run on NVIDIA GPUs, compiling …
632137active
soedinglab/MMseqs2
MMseqs2 is an open-source C++ suite for ultra-fast, sensitive search and clustering of huge protein and nucleotide sequence sets, running o…
702131active
deepspeedai/DeepSpeed-MII
DeepSpeed-MII is a Python library for high-throughput, low-latency large language model inference, built on DeepSpeed. It provides blocked …
432111active
inducer/pycuda
PyCUDA is a Python wrapper giving Pythonic access to Nvidia's CUDA parallel computation API, including GPUArray multidimensional arrays and…
782051active
lightseekorg/tokenspeed
TokenSpeed is a high-performance LLM inference engine designed for agentic workloads, aiming for TensorRT-LLM-level performance with vLLM-l…
681989active
NVIDIA/Stable-Diffusion-WebUI-TensorRT
An NVIDIA extension for the Stable Diffusion Web UI (Automatic1111) that accelerates image generation using TensorRT-optimized engines on R…
181989active
develsoftware/GMinerRelease
GMiner is a closed-source GPU cryptocurrency miner for NVIDIA and AMD cards supporting algorithms such as Ethash, ProgPoW, KAWPOW, Equihash…
671981active
siliconflow/onediff
OneDiff is an out-of-the-box acceleration library for diffusion models, providing PyTorch compilation tools and optimized GPU kernels. It i…
481964active
openmm/openmm
OpenMM is a high-performance toolkit and library for molecular dynamics simulation, with optimized GPU-accelerated kernels. It can be used …
951960stable
CannyLab/tsne-cuda
A CUDA-accelerated implementation of the FIt-SNE t-SNE algorithm with Python bindings, offering up to 1200x speedup over scikit-learn. It e…
951957stable
AdaptiveCpp/AdaptiveCpp
AdaptiveCpp (formerly hipSYCL/Open SYCL) is an independent, community-driven C++ compiler platform for heterogeneous programming models inc…
751928active
rednote-machine-learning/RedKnot
RedKnot is a long-context LLM inference acceleration library built on SGLang, using head-classified KV reuse, offline KV storage with RoPE …
581880active
laekov/fastmoe
FastMoE is a PyTorch library providing efficient Mixture of Experts (MoE) layers with custom C/CUDA operators. It supports distributed expe…
261859active
dotnet/TorchSharp
TorchSharp is a .NET library providing bindings to LibTorch, the library that powers PyTorch, with a focus on tensors and a PyTorch-like AP…
731850active
dphnAI/sonar
Sonar is a large-scale LLM inference engine for Hugging Face-compatible language and multimodal models, based on vLLM (formerly the Aphrodi…
941843active
zju3dv/4K4D
4K4D is a research implementation of a 4D point cloud representation for real-time dynamic view synthesis at up to 4K resolution, built on …
271807active
FL33TW00D/whisper-turbo
Whisper Turbo is a fast, cross-platform, GPU-accelerated implementation of OpenAI's Whisper speech recognition model, built on the Ratchet …
191794active
moderngpu/moderngpu
moderngpu is a header-only C++ productivity library for general-purpose GPU computing built on CUDA. It provides accelerated primitives for…
511789active
su2code/SU2
SU2 is an open-source C++ suite for numerically solving partial differential equations and performing PDE-constrained optimization, primari…
841784active
chenjd/Render-Crowd-Of-Animated-Characters
A Unity library that bakes skeletal animation data into an animation map texture and uses GPU instancing to render tens of thousands of ani…
321774active
DTolm/VkFFT
VkFFT is an open-source, GPU-accelerated multidimensional Fast Fourier Transform library supporting Vulkan, CUDA, HIP, OpenCL, Level Zero, …
571769active
Coyote-A/ultimate-upscale-for-automatic1111
An extension for the AUTOMATIC1111 Stable Diffusion web UI that upscales images to 2K/4K+ by processing them in tiled passes with diffusion…
311766stable
webonnx/wonnx
Wonnx is a GPU-accelerated ONNX inference runtime written entirely in Rust, built on wgpu and usable natively or in the browser via WebGPU …
101754active
m4rs-mt/ILGPU
ILGPU is a just-in-time compiler for high-performance GPU programs (kernels) written in .NET languages like C# and F#, entirely in C# with …
681748active
rusty1s/pytorch_scatter
A PyTorch extension library providing highly optimized scatter and segment (sparse update) operations with sum, mean, min, and max reductio…
611745stable
deepseek-ai/TileKernels
TileKernels is a Python library of optimized GPU kernels for LLM operations, written in the TileLang DSL. It provides kernels for MoE routi…
491743active
0xSero/turboquant
TurboQuant is a Python library implementing near-optimal KV cache quantization for LLM inference, compressing keys to 3-bit and values to 2…
481739active
tile-ai/TileRT
TileRT is a tile-based runtime for ultra-low-latency LLM inference, achieving hundreds to over 1000 tokens/s decode speeds on frontier mode…
771738active
Zaneham/Booth
Booth is an open-source compiler that takes CUDA C, HIP, or Triton kernel source and emits binaries for AMD RDNA 2/3/4 GPUs, NVIDIA PTX, Te…
771731active
jiaweizzhao/GaLore
GaLore is a PyTorch library implementing Gradient Low-Rank Projection for memory-efficient full-parameter training of large language models…
251700active
EnzymeAD/Enzyme
Enzyme is a high-performance automatic differentiation plugin for LLVM and MLIR that computes derivatives and gradients of arbitrary existi…
951679active
NVIDIA/nccl-tests
NVIDIA's official test suite for NCCL that measures the performance and correctness of collective communication operations across GPUs. It …
761636active
chainer/chainer
Chainer is a Python-based deep learning framework that pioneered the define-by-run approach with dynamic computational graphs and automatic…
235924maintenance
ethereum-mining/ethminer
Ethminer is a command-line GPU mining application for Ethash Proof of Work coins such as Ethereum and Ethereum Classic. It supports OpenCL …
105918maintenance
databricks/megablocks
MegaBlocks is a lightweight Python library for efficient training of mixture-of-experts (MoE) models, built around its dropless-MoE (dMoE) …
621588active
xLLM-AI/xllm
xLLM is a high-performance C++ inference engine for LLM, VLM, DiT and recommendation models, optimized for heterogeneous AI accelerators su…
781536active
RightNow-AI/autokernel
AutoKernel is an open-source autoresearch pipeline that takes any PyTorch model, profiles it to find GPU kernel bottlenecks, extracts them …
481535active
leela-zero/leela-zero
Leela Zero is an open-source Go engine that reimplements AlphaGo Zero, combining Monte Carlo Tree Search with a deep residual convolutional…
235586maintenance
ByteDance-Seed/Triton-distributed
Triton-distributed is a distributed compiler built on OpenAI Triton for computation-communication overlapping on multi-GPU systems. It lets…
631526active
intel/llvm
Intel's staging area for LLVM upstream contributions and home for Intel LLVM-based projects, most notably the oneAPI DPC++ compiler impleme…
951517active
uccl-project/uccl
UCCL is a high-performance GPU communication library written in C++ that provides collectives (as a drop-in NCCL/RCCL replacement), P2P tra…
711496active
beehive-lab/TornadoVM
TornadoVM is a Java programming framework that JIT-compiles JVM bytecode into GPU kernels targeting NVIDIA CUDA, OpenCL, and Apple Metal at…
971491active
graphdeco-inria/diff-gaussian-rasterization
A CUDA-based differentiable rasterization engine for 3D Gaussian Splatting, used in the SIGGRAPH 2023 paper '3D Gaussian Splatting for Real…
291489stable
mit-han-lab/torchsparse
TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a…
261472active
jax-md/jax-md
JAX MD is a Python library for molecular dynamics simulations built on JAX, making them hardware accelerated on CPU, GPU, and TPU and end-t…
861458active
NVIDIA/MatX
MatX is a C++20 header-only numerical computing library providing NumPy-like tensor expressions for NVIDIA GPUs and multithreaded CPUs. It …
901444active
intel/compute-runtime
Intel's open-source Graphics Compute Runtime (NEO) providing OpenCL 3.0 and oneAPI Level Zero compute API support for Intel HD Graphics and…
951432active
JuliaGPU/CUDA.jl
CUDA.jl is the main Julia package for programming NVIDIA CUDA GPUs, offering a high-level CuArray abstraction, a compiler for writing CUDA …
991424stable
CliMA/Oceananigans.jl
Oceananigans.jl is a Julia package for fast, flexible finite-volume simulations of incompressible fluid dynamics, solving nonhydrostatic an…
951412active
NVIDIA/gdrcopy
GDRCopy is a low-latency GPU memory copy library built on NVIDIA GPUDirect RDMA technology, allowing CPU-driven copies between host and GPU…
841412active
heterodb/pg-strom
PG-Strom is a PostgreSQL extension that accelerates SQL analytics and batch workloads using GPU devices, NVMe-SSD storage, and Apache Arrow…
761408active
RahulSChand/gpu_poor
A web-based calculator that estimates GPU memory requirements and inference/finetuning throughput (token/s) for any LLM. It supports quanti…
181405active
NVIDIA/thrust
Thrust is a C++ parallel algorithms library providing an STL-like high-level interface for GPU and multicore CPU computing, built on CUDA, …
105002maintenance
uxlfoundation/scikit-learn-intelex
Intel's Extension for Scikit-learn is a free AI accelerator that speeds up existing scikit-learn workflows on CPUs and GPUs, claiming up to…
911356active
bytedance/flux
Flux is a GPU kernel library from ByteDance that overlaps computation with communication for tensor and expert parallelism in dense and MoE…
331354active
hao-ai-lab/LookaheadDecoding
A Python library implementing Lookahead Decoding, an exact parallel decoding algorithm that accelerates LLM inference without a draft model…
311342active
mapillary/inplace_abn
A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing…
651333stable
LuxCoreRender/LuxCore
LuxCoreRender is a physically based, unbiased rendering engine with a C++ and Python API (LuxCore), supporting CPU, OpenCL, CUDA, and OptiX…
921321active
alibaba/rtp-llm
RTP-LLM is Alibaba's high-performance LLM inference engine written in C++/CUDA, optimized with kernels like PagedAttention and FlashAttenti…
661316active
ChenmienTan/RL2
RL2 (Ray Less Reinforcement Learning) is a concise Python library for post-training large language models with reinforcement learning, SFT,…
571307active
steineggerlab/foldseek
Foldseek is a command-line tool for fast and sensitive comparison, search, and clustering of large protein 3D structure sets, including mon…
671283active
turboderp-org/exllamav3
ExLlamaV3 is a Python library for fast quantization and inference of large language models on consumer-class GPUs, featuring the EXL3 quant…
871274active
TheRock
TheRock is AMD's open-source build and release system for the ROCm software stack, replacing the legacy monolithic ROCm release process wit…
891270active
stotko/stdgpu
stdgpu is a lightweight C++17 library providing STL-like generic data structures (vector, unordered_map, unordered_set, deque, queue, stack…
641270active
vipshop/cache-dit
Cache-DiT is a PyTorch-native inference engine that accelerates Diffusion Transformer (DiT) models with hybrid caching, parallelism, quanti…
821267active
corsix/amx
A C library and documentation of Apple's undocumented AMX (Apple Matrix eXtension) instructions found on M1-M4 Apple Silicon chips, enablin…
321261active
MoonshotAI/FlashKDA
FlashKDA is a set of high-performance CUDA kernels (built on CUTLASS) implementing Kimi Delta Attention, a linear attention mechanism, for …
571229active
microsoft/MInference
MInference is a Microsoft library that accelerates long-context LLM inference using dynamic sparse attention, reducing pre-fill latency by …
491226active
wcandillon/react-native-webgpu
A React Native library that brings the WebGPU API to mobile and desktop apps, powered by Dawn (Chrome's WebGPU implementation). It supports…
911225active
eduardoleao052/js-pytorch
JS-PyTorch is a deep learning library for JavaScript that closely mirrors PyTorch's syntax, providing tensor operations, automatic differen…
161222active
chelsea0x3b/cudarc
cudarc is a safe, minimal Rust wrapper around the NVIDIA CUDA toolkit, exposing the CUDA driver API plus libraries such as NVRTC, cuBLAS/cu…
981213active
cp2k/cp2k
CP2K is an open-source quantum chemistry and solid state physics package for atomistic simulations of molecular, liquid, periodic, and biol…
891198stable
getkeops/keops
KeOps (pykeops) is a Python library for computing kernel reductions over large arrays on CPUs and GPUs using efficient C++/CUDA routines wi…
651189active
CNugteren/CLBlast
CLBlast is a lightweight, tunable OpenCL BLAS library written in C++11 that implements basic linear algebra subprograms for vectors and mat…
701186stable
higgsfield-ai/higgsfield
Higgsfield is an open-source GPU orchestration and machine learning framework for fault-tolerant, distributed training of very large models…
234106maintenance
AlgRUC/JittorGeometric
JittorGeometric is a graph machine learning library built on the Jittor deep learning framework, providing implementations of 40+ Graph Neu…
621177active
keijiro/StableFluids
A GPU-based Unity implementation of Jos Stam's Stable Fluids fluid simulation using compute shaders. It exposes the velocity field for rend…
501172active
baidu-research/warp-ctc
A fast parallel implementation of the Connectionist Temporal Classification (CTC) loss function for CPU and CUDA GPU, with a simple C inter…
324069maintenance
XMR-Stak
xmr-stak is a free, open-source, high-performance miner for Monero (RandomX) and unified CryptoNight-based cryptocurrencies, supporting CPU…
234059maintenance
inducer/pyopencl
PyOpenCL is a Python wrapper providing Pythonic access to the OpenCL parallel computation API, letting you run kernels on GPUs and other ma…
991149stable
sgl-project/SpecForge
SpecForge is a Python framework from the SGLang team for training speculative decoding models such as EAGLE/EAGLE3 draft heads. Trained mod…
641145active
extropic-ai/thrml
THRML is a JAX library for building and sampling probabilistic graphical models, focused on efficient block Gibbs sampling of energy-based …
711144active
ovg-project/kvcached
kvcached is a Python library that brings OS-style virtual memory abstraction to KV cache management for LLM serving and training on shared …
741143active
nnaisense/evotorch
EvoTorch is an open-source evolutionary computation library built on top of PyTorch, developed at NNAISENSE. It provides distribution-based…
781142active
IST-DASLab/marlin
Marlin is a highly optimized FP16xINT4 matrix multiplication CUDA kernel for LLM inference that achieves near-ideal 4x speedups at batch si…
261136active
Dao-AILab/quack
QuACK is a collection of high-performance GPU kernels (RMSNorm, LayerNorm, softmax, cross-entropy, GEMM with epilogues) written in NVIDIA's…
851135active
Tencent/hpc-ops
HPC-Ops is a production-grade C++/CUDA operator library for high-performance LLM inference, developed by Tencent's Hunyuan AI Infra team. I…
591131active
uncomplicate/neanderthal
Neanderthal is a fast Clojure library for matrix and linear algebra computations built on optimized native BLAS and LAPACK routines, suppor…
761128active
googlecolab/google-colab-cli
A Python-based command-line interface for Google Colab that lets users provision CPU, GPU, and TPU runtimes, execute local scripts and note…
591126active
Lasagne/Lasagne
Lasagne is a lightweight Python library for building and training neural networks on top of Theano. It supports feed-forward, convolutional…
233857maintenance
rusty1s/pytorch_sparse
A PyTorch extension library providing optimized sparse matrix operations (coalesce, transpose, sparse-dense and sparse-sparse multiplicatio…
611104active
MoonshotAI/MoonEP
MoonEP is an Expert Parallelism communication library for Mixture-of-Experts training that keeps token loads perfectly balanced across rank…
561101active
luchris429/purejaxrl
PureJaxRL provides end-to-end reinforcement learning training pipelines implemented entirely in JAX, including environments, enabling massi…
301099active
pocl/pocl
PoCL (Portable Computing Language) is an MIT-licensed, conformant open-source implementation of the OpenCL 3.0 standard that can be easily …
741075active
NVIDIA-Merlin/HugeCTR
HugeCTR is a GPU-accelerated deep learning framework from NVIDIA designed for training and inference of large recommender models, especiall…
771071active
sirius-db/sirius
Sirius is a GPU-native SQL analytics engine written in C++ that accelerates query execution by offloading it to GPUs. It integrates with ex…
691059active

← prev page 2 / 6 next →