function: gpu-computing
559 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| wilicc/gpu-burn GPU Burn is a multi-GPU CUDA stress test tool that pushes NVIDIA GPUs to maximum load for stability and thermal testing. It supports config… | 70 | 2322 | active |
| apache/mahout Apache Mahout is an Apache Software Foundation project providing Qumat, a high-level Python library for quantum computing with a unified AP… | 86 | 2304 | active |
| CoinCheung/pytorch-loss A PyTorch library providing a collection of loss functions (focal loss, triplet loss, AMSoftmax, label-smooth CE, dice loss, lovasz-softmax… | 32 | 2252 | active |
| kubeflow/trainer Kubeflow Trainer is a Kubernetes-native platform for distributed AI model training and LLM fine-tuning across frameworks like PyTorch, JAX,… | 95 | 2198 | active |
| NVIDIA/cutile-python cuTile Python is a tile-based programming language and library from NVIDIA for writing parallel kernels that run on NVIDIA GPUs, compiling … | 63 | 2137 | active |
| soedinglab/MMseqs2 MMseqs2 is an open-source C++ suite for ultra-fast, sensitive search and clustering of huge protein and nucleotide sequence sets, running o… | 70 | 2131 | active |
| deepspeedai/DeepSpeed-MII DeepSpeed-MII is a Python library for high-throughput, low-latency large language model inference, built on DeepSpeed. It provides blocked … | 43 | 2111 | active |
| inducer/pycuda PyCUDA is a Python wrapper giving Pythonic access to Nvidia's CUDA parallel computation API, including GPUArray multidimensional arrays and… | 78 | 2051 | active |
| lightseekorg/tokenspeed TokenSpeed is a high-performance LLM inference engine designed for agentic workloads, aiming for TensorRT-LLM-level performance with vLLM-l… | 68 | 1989 | active |
| NVIDIA/Stable-Diffusion-WebUI-TensorRT An NVIDIA extension for the Stable Diffusion Web UI (Automatic1111) that accelerates image generation using TensorRT-optimized engines on R… | 18 | 1989 | active |
| develsoftware/GMinerRelease GMiner is a closed-source GPU cryptocurrency miner for NVIDIA and AMD cards supporting algorithms such as Ethash, ProgPoW, KAWPOW, Equihash… | 67 | 1981 | active |
| siliconflow/onediff OneDiff is an out-of-the-box acceleration library for diffusion models, providing PyTorch compilation tools and optimized GPU kernels. It i… | 48 | 1964 | active |
| openmm/openmm OpenMM is a high-performance toolkit and library for molecular dynamics simulation, with optimized GPU-accelerated kernels. It can be used … | 95 | 1960 | stable |
| CannyLab/tsne-cuda A CUDA-accelerated implementation of the FIt-SNE t-SNE algorithm with Python bindings, offering up to 1200x speedup over scikit-learn. It e… | 95 | 1957 | stable |
| AdaptiveCpp/AdaptiveCpp AdaptiveCpp (formerly hipSYCL/Open SYCL) is an independent, community-driven C++ compiler platform for heterogeneous programming models inc… | 75 | 1928 | active |
| rednote-machine-learning/RedKnot RedKnot is a long-context LLM inference acceleration library built on SGLang, using head-classified KV reuse, offline KV storage with RoPE … | 58 | 1880 | active |
| laekov/fastmoe FastMoE is a PyTorch library providing efficient Mixture of Experts (MoE) layers with custom C/CUDA operators. It supports distributed expe… | 26 | 1859 | active |
| dotnet/TorchSharp TorchSharp is a .NET library providing bindings to LibTorch, the library that powers PyTorch, with a focus on tensors and a PyTorch-like AP… | 73 | 1850 | active |
| dphnAI/sonar Sonar is a large-scale LLM inference engine for Hugging Face-compatible language and multimodal models, based on vLLM (formerly the Aphrodi… | 94 | 1843 | active |
| zju3dv/4K4D 4K4D is a research implementation of a 4D point cloud representation for real-time dynamic view synthesis at up to 4K resolution, built on … | 27 | 1807 | active |
| FL33TW00D/whisper-turbo Whisper Turbo is a fast, cross-platform, GPU-accelerated implementation of OpenAI's Whisper speech recognition model, built on the Ratchet … | 19 | 1794 | active |
| moderngpu/moderngpu moderngpu is a header-only C++ productivity library for general-purpose GPU computing built on CUDA. It provides accelerated primitives for… | 51 | 1789 | active |
| su2code/SU2 SU2 is an open-source C++ suite for numerically solving partial differential equations and performing PDE-constrained optimization, primari… | 84 | 1784 | active |
| chenjd/Render-Crowd-Of-Animated-Characters A Unity library that bakes skeletal animation data into an animation map texture and uses GPU instancing to render tens of thousands of ani… | 32 | 1774 | active |
| DTolm/VkFFT VkFFT is an open-source, GPU-accelerated multidimensional Fast Fourier Transform library supporting Vulkan, CUDA, HIP, OpenCL, Level Zero, … | 57 | 1769 | active |
| Coyote-A/ultimate-upscale-for-automatic1111 An extension for the AUTOMATIC1111 Stable Diffusion web UI that upscales images to 2K/4K+ by processing them in tiled passes with diffusion… | 31 | 1766 | stable |
| webonnx/wonnx Wonnx is a GPU-accelerated ONNX inference runtime written entirely in Rust, built on wgpu and usable natively or in the browser via WebGPU … | 10 | 1754 | active |
| m4rs-mt/ILGPU ILGPU is a just-in-time compiler for high-performance GPU programs (kernels) written in .NET languages like C# and F#, entirely in C# with … | 68 | 1748 | active |
| rusty1s/pytorch_scatter A PyTorch extension library providing highly optimized scatter and segment (sparse update) operations with sum, mean, min, and max reductio… | 61 | 1745 | stable |
| deepseek-ai/TileKernels TileKernels is a Python library of optimized GPU kernels for LLM operations, written in the TileLang DSL. It provides kernels for MoE routi… | 49 | 1743 | active |
| 0xSero/turboquant TurboQuant is a Python library implementing near-optimal KV cache quantization for LLM inference, compressing keys to 3-bit and values to 2… | 48 | 1739 | active |
| tile-ai/TileRT TileRT is a tile-based runtime for ultra-low-latency LLM inference, achieving hundreds to over 1000 tokens/s decode speeds on frontier mode… | 77 | 1738 | active |
| Zaneham/Booth Booth is an open-source compiler that takes CUDA C, HIP, or Triton kernel source and emits binaries for AMD RDNA 2/3/4 GPUs, NVIDIA PTX, Te… | 77 | 1731 | active |
| jiaweizzhao/GaLore GaLore is a PyTorch library implementing Gradient Low-Rank Projection for memory-efficient full-parameter training of large language models… | 25 | 1700 | active |
| EnzymeAD/Enzyme Enzyme is a high-performance automatic differentiation plugin for LLVM and MLIR that computes derivatives and gradients of arbitrary existi… | 95 | 1679 | active |
| NVIDIA/nccl-tests NVIDIA's official test suite for NCCL that measures the performance and correctness of collective communication operations across GPUs. It … | 76 | 1636 | active |
| chainer/chainer Chainer is a Python-based deep learning framework that pioneered the define-by-run approach with dynamic computational graphs and automatic… | 23 | 5924 | maintenance |
| ethereum-mining/ethminer Ethminer is a command-line GPU mining application for Ethash Proof of Work coins such as Ethereum and Ethereum Classic. It supports OpenCL … | 10 | 5918 | maintenance |
| databricks/megablocks MegaBlocks is a lightweight Python library for efficient training of mixture-of-experts (MoE) models, built around its dropless-MoE (dMoE) … | 62 | 1588 | active |
| xLLM-AI/xllm xLLM is a high-performance C++ inference engine for LLM, VLM, DiT and recommendation models, optimized for heterogeneous AI accelerators su… | 78 | 1536 | active |
| RightNow-AI/autokernel AutoKernel is an open-source autoresearch pipeline that takes any PyTorch model, profiles it to find GPU kernel bottlenecks, extracts them … | 48 | 1535 | active |
| leela-zero/leela-zero Leela Zero is an open-source Go engine that reimplements AlphaGo Zero, combining Monte Carlo Tree Search with a deep residual convolutional… | 23 | 5586 | maintenance |
| ByteDance-Seed/Triton-distributed Triton-distributed is a distributed compiler built on OpenAI Triton for computation-communication overlapping on multi-GPU systems. It lets… | 63 | 1526 | active |
| intel/llvm Intel's staging area for LLVM upstream contributions and home for Intel LLVM-based projects, most notably the oneAPI DPC++ compiler impleme… | 95 | 1517 | active |
| uccl-project/uccl UCCL is a high-performance GPU communication library written in C++ that provides collectives (as a drop-in NCCL/RCCL replacement), P2P tra… | 71 | 1496 | active |
| beehive-lab/TornadoVM TornadoVM is a Java programming framework that JIT-compiles JVM bytecode into GPU kernels targeting NVIDIA CUDA, OpenCL, and Apple Metal at… | 97 | 1491 | active |
| graphdeco-inria/diff-gaussian-rasterization A CUDA-based differentiable rasterization engine for 3D Gaussian Splatting, used in the SIGGRAPH 2023 paper '3D Gaussian Splatting for Real… | 29 | 1489 | stable |
| mit-han-lab/torchsparse TorchSparse is a high-performance PyTorch library for sparse convolution on 3D point clouds, with optimized GPU kernels for both training a… | 26 | 1472 | active |
| jax-md/jax-md JAX MD is a Python library for molecular dynamics simulations built on JAX, making them hardware accelerated on CPU, GPU, and TPU and end-t… | 86 | 1458 | active |
| NVIDIA/MatX MatX is a C++20 header-only numerical computing library providing NumPy-like tensor expressions for NVIDIA GPUs and multithreaded CPUs. It … | 90 | 1444 | active |
| intel/compute-runtime Intel's open-source Graphics Compute Runtime (NEO) providing OpenCL 3.0 and oneAPI Level Zero compute API support for Intel HD Graphics and… | 95 | 1432 | active |
| JuliaGPU/CUDA.jl CUDA.jl is the main Julia package for programming NVIDIA CUDA GPUs, offering a high-level CuArray abstraction, a compiler for writing CUDA … | 99 | 1424 | stable |
| CliMA/Oceananigans.jl Oceananigans.jl is a Julia package for fast, flexible finite-volume simulations of incompressible fluid dynamics, solving nonhydrostatic an… | 95 | 1412 | active |
| NVIDIA/gdrcopy GDRCopy is a low-latency GPU memory copy library built on NVIDIA GPUDirect RDMA technology, allowing CPU-driven copies between host and GPU… | 84 | 1412 | active |
| heterodb/pg-strom PG-Strom is a PostgreSQL extension that accelerates SQL analytics and batch workloads using GPU devices, NVMe-SSD storage, and Apache Arrow… | 76 | 1408 | active |
| RahulSChand/gpu_poor A web-based calculator that estimates GPU memory requirements and inference/finetuning throughput (token/s) for any LLM. It supports quanti… | 18 | 1405 | active |
| NVIDIA/thrust Thrust is a C++ parallel algorithms library providing an STL-like high-level interface for GPU and multicore CPU computing, built on CUDA, … | 10 | 5002 | maintenance |
| uxlfoundation/scikit-learn-intelex Intel's Extension for Scikit-learn is a free AI accelerator that speeds up existing scikit-learn workflows on CPUs and GPUs, claiming up to… | 91 | 1356 | active |
| bytedance/flux Flux is a GPU kernel library from ByteDance that overlaps computation with communication for tensor and expert parallelism in dense and MoE… | 33 | 1354 | active |
| hao-ai-lab/LookaheadDecoding A Python library implementing Lookahead Decoding, an exact parallel decoding algorithm that accelerates LLM inference without a draft model… | 31 | 1342 | active |
| mapillary/inplace_abn A PyTorch extension library implementing In-Place Activated BatchNorm (InPlace-ABN), which redefines BN plus nonlinear activation as a sing… | 65 | 1333 | stable |
| LuxCoreRender/LuxCore LuxCoreRender is a physically based, unbiased rendering engine with a C++ and Python API (LuxCore), supporting CPU, OpenCL, CUDA, and OptiX… | 92 | 1321 | active |
| alibaba/rtp-llm RTP-LLM is Alibaba's high-performance LLM inference engine written in C++/CUDA, optimized with kernels like PagedAttention and FlashAttenti… | 66 | 1316 | active |
| ChenmienTan/RL2 RL2 (Ray Less Reinforcement Learning) is a concise Python library for post-training large language models with reinforcement learning, SFT,… | 57 | 1307 | active |
| steineggerlab/foldseek Foldseek is a command-line tool for fast and sensitive comparison, search, and clustering of large protein 3D structure sets, including mon… | 67 | 1283 | active |
| turboderp-org/exllamav3 ExLlamaV3 is a Python library for fast quantization and inference of large language models on consumer-class GPUs, featuring the EXL3 quant… | 87 | 1274 | active |
| TheRock TheRock is AMD's open-source build and release system for the ROCm software stack, replacing the legacy monolithic ROCm release process wit… | 89 | 1270 | active |
| stotko/stdgpu stdgpu is a lightweight C++17 library providing STL-like generic data structures (vector, unordered_map, unordered_set, deque, queue, stack… | 64 | 1270 | active |
| vipshop/cache-dit Cache-DiT is a PyTorch-native inference engine that accelerates Diffusion Transformer (DiT) models with hybrid caching, parallelism, quanti… | 82 | 1267 | active |
| corsix/amx A C library and documentation of Apple's undocumented AMX (Apple Matrix eXtension) instructions found on M1-M4 Apple Silicon chips, enablin… | 32 | 1261 | active |
| MoonshotAI/FlashKDA FlashKDA is a set of high-performance CUDA kernels (built on CUTLASS) implementing Kimi Delta Attention, a linear attention mechanism, for … | 57 | 1229 | active |
| microsoft/MInference MInference is a Microsoft library that accelerates long-context LLM inference using dynamic sparse attention, reducing pre-fill latency by … | 49 | 1226 | active |
| wcandillon/react-native-webgpu A React Native library that brings the WebGPU API to mobile and desktop apps, powered by Dawn (Chrome's WebGPU implementation). It supports… | 91 | 1225 | active |
| eduardoleao052/js-pytorch JS-PyTorch is a deep learning library for JavaScript that closely mirrors PyTorch's syntax, providing tensor operations, automatic differen… | 16 | 1222 | active |
| chelsea0x3b/cudarc cudarc is a safe, minimal Rust wrapper around the NVIDIA CUDA toolkit, exposing the CUDA driver API plus libraries such as NVRTC, cuBLAS/cu… | 98 | 1213 | active |
| cp2k/cp2k CP2K is an open-source quantum chemistry and solid state physics package for atomistic simulations of molecular, liquid, periodic, and biol… | 89 | 1198 | stable |
| getkeops/keops KeOps (pykeops) is a Python library for computing kernel reductions over large arrays on CPUs and GPUs using efficient C++/CUDA routines wi… | 65 | 1189 | active |
| CNugteren/CLBlast CLBlast is a lightweight, tunable OpenCL BLAS library written in C++11 that implements basic linear algebra subprograms for vectors and mat… | 70 | 1186 | stable |
| higgsfield-ai/higgsfield Higgsfield is an open-source GPU orchestration and machine learning framework for fault-tolerant, distributed training of very large models… | 23 | 4106 | maintenance |
| AlgRUC/JittorGeometric JittorGeometric is a graph machine learning library built on the Jittor deep learning framework, providing implementations of 40+ Graph Neu… | 62 | 1177 | active |
| keijiro/StableFluids A GPU-based Unity implementation of Jos Stam's Stable Fluids fluid simulation using compute shaders. It exposes the velocity field for rend… | 50 | 1172 | active |
| baidu-research/warp-ctc A fast parallel implementation of the Connectionist Temporal Classification (CTC) loss function for CPU and CUDA GPU, with a simple C inter… | 32 | 4069 | maintenance |
| XMR-Stak xmr-stak is a free, open-source, high-performance miner for Monero (RandomX) and unified CryptoNight-based cryptocurrencies, supporting CPU… | 23 | 4059 | maintenance |
| inducer/pyopencl PyOpenCL is a Python wrapper providing Pythonic access to the OpenCL parallel computation API, letting you run kernels on GPUs and other ma… | 99 | 1149 | stable |
| sgl-project/SpecForge SpecForge is a Python framework from the SGLang team for training speculative decoding models such as EAGLE/EAGLE3 draft heads. Trained mod… | 64 | 1145 | active |
| extropic-ai/thrml THRML is a JAX library for building and sampling probabilistic graphical models, focused on efficient block Gibbs sampling of energy-based … | 71 | 1144 | active |
| ovg-project/kvcached kvcached is a Python library that brings OS-style virtual memory abstraction to KV cache management for LLM serving and training on shared … | 74 | 1143 | active |
| nnaisense/evotorch EvoTorch is an open-source evolutionary computation library built on top of PyTorch, developed at NNAISENSE. It provides distribution-based… | 78 | 1142 | active |
| IST-DASLab/marlin Marlin is a highly optimized FP16xINT4 matrix multiplication CUDA kernel for LLM inference that achieves near-ideal 4x speedups at batch si… | 26 | 1136 | active |
| Dao-AILab/quack QuACK is a collection of high-performance GPU kernels (RMSNorm, LayerNorm, softmax, cross-entropy, GEMM with epilogues) written in NVIDIA's… | 85 | 1135 | active |
| Tencent/hpc-ops HPC-Ops is a production-grade C++/CUDA operator library for high-performance LLM inference, developed by Tencent's Hunyuan AI Infra team. I… | 59 | 1131 | active |
| uncomplicate/neanderthal Neanderthal is a fast Clojure library for matrix and linear algebra computations built on optimized native BLAS and LAPACK routines, suppor… | 76 | 1128 | active |
| googlecolab/google-colab-cli A Python-based command-line interface for Google Colab that lets users provision CPU, GPU, and TPU runtimes, execute local scripts and note… | 59 | 1126 | active |
| Lasagne/Lasagne Lasagne is a lightweight Python library for building and training neural networks on top of Theano. It supports feed-forward, convolutional… | 23 | 3857 | maintenance |
| rusty1s/pytorch_sparse A PyTorch extension library providing optimized sparse matrix operations (coalesce, transpose, sparse-dense and sparse-sparse multiplicatio… | 61 | 1104 | active |
| MoonshotAI/MoonEP MoonEP is an Expert Parallelism communication library for Mixture-of-Experts training that keeps token loads perfectly balanced across rank… | 56 | 1101 | active |
| luchris429/purejaxrl PureJaxRL provides end-to-end reinforcement learning training pipelines implemented entirely in JAX, including environments, enabling massi… | 30 | 1099 | active |
| pocl/pocl PoCL (Portable Computing Language) is an MIT-licensed, conformant open-source implementation of the OpenCL 3.0 standard that can be easily … | 74 | 1075 | active |
| NVIDIA-Merlin/HugeCTR HugeCTR is a GPU-accelerated deep learning framework from NVIDIA designed for training and inference of large recommender models, especiall… | 77 | 1071 | active |
| sirius-db/sirius Sirius is a GPU-native SQL analytics engine written in C++ that accelerates query execution by offloading it to GPUs. It integrates with ex… | 69 | 1059 | active |