function: gpu-computing
559 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| abiosoft/colima Colima is a CLI tool that provides container runtimes (Docker, Containerd, Incus) on macOS and Linux with minimal setup, built on Lima. It … | 95 | 30526 | active |
| modular/modular Modular Platform hosting the MAX AI serving framework and the Mojo systems programming language. It provides an OpenAI-compatible inference… | 92 | 29225 | active |
| MLX MLX is an array computation framework for machine learning on Apple silicon, developed by Apple ML research. It offers NumPy-like Python AP… | 94 | 28172 | active |
| Stability-AI/generative-models Stability AI's official repository of generative diffusion models, including Stable Diffusion, Stable Video, and SV4D 2.0 for image, video,… | 45 | 27271 | active |
| ApolloAuto/apollo Apollo is an open-source autonomous driving platform providing a high-performance, modular software stack for developing, testing, and depl… | 57 | 26807 | active |
| PaddlePaddle/Paddle PaddlePaddle is an industrial-grade deep learning framework written in C++ with Python APIs, supporting high-performance single-machine and… | 84 | 24062 | active |
| Tencent/ncnn ncnn is a high-performance neural network inference framework written in C++ and optimized for mobile, embedded, and desktop deployment. It… | 86 | 23753 | stable |
| verl-project/verl verl (Volcano Engine Reinforcement Learning) is a flexible, production-ready RL post-training library for large language models, open-sourc… | 84 | 23145 | active |
| microsoft/onnxruntime ONNX Runtime is a cross-platform, high-performance machine-learning accelerator for running inference and training on ONNX models. It suppo… | 99 | 21654 | stable |
| k4yt3x/video2x Video2X is a machine learning-based video super-resolution and frame interpolation framework written in C/C++. It upscales videos and image… | 66 | 21263 | active |
| huggingface/candle Candle is a minimalist machine learning framework for Rust focused on performance and ease of use, with CPU and CUDA GPU support. It ships … | 73 | 20955 | active |
| NVIDIA/Megatron-LM NVIDIA's GPU-optimized library for training large transformer models at scale, comprising Megatron-LM (reference training scripts) and Mega… | 99 | 17615 | active |
| alibaba/MNN MNN is a lightweight, high-performance deep learning inference engine developed by Alibaba, supporting LLMs, vision, audio, and multimodal … | 93 | 15973 | active |
| tracel-ai/burn Burn is a Rust-based tensor library and deep learning framework supporting training and inference through a unified API. It JIT-compiles te… | 89 | 15816 | active |
| GeeeekExplorer/nano-vllm A lightweight vLLM-style LLM inference engine implemented from scratch in about 1,200 lines of Python. It offers fast offline inference wit… | 54 | 15164 | active |
| facebookresearch/vggt VGGT (Visual Geometry Grounded Transformer) is a feed-forward transformer model from Meta AI and Oxford VGG that infers 3D geometry—camera … | 58 | 14292 | active |
| Eclipse Deeplearning4J Eclipse Deeplearning4J is an open-source deep learning framework and ecosystem for the JVM, including the ND4J linear algebra library, the … | 77 | 14246 | active |
| Open3D Open3D is an open-source C++ and Python library for 3D data processing, offering data structures, algorithms, and pipelines for point cloud… | 67 | 13913 | active |
| apache/tvm Apache TVM is an open machine learning compiler framework that takes pre-trained models and compiles them into optimized, deployable module… | 90 | 13691 | active |
| openwall/john John the Ripper jumbo is an open-source offline password cracker supporting hundreds of hash and cipher types, from Unix and Windows passwo… | 75 | 13543 | active |
| NVIDIA/TensorRT NVIDIA TensorRT is an SDK for high-performance deep learning inference on NVIDIA GPUs, comprising an inference compiler that converts train… | 94 | 13293 | stable |
| modelscope/DiffSynth-Studio DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio … | 79 | 13003 | active |
| cupy/cupy CuPy is a NumPy/SciPy-compatible array library for GPU-accelerated computing with Python, running on NVIDIA CUDA or AMD ROCm. It acts as a … | 95 | 12278 | stable |
| LMCache/LMCache LMCache is a KV cache management layer for LLM inference that stores, compresses, and reuses KV caches across requests, sessions, and servi… | 86 | 11460 | active |
| triton-inference-server/server NVIDIA Triton Inference Server is an open-source inference serving software that deploys AI models from multiple frameworks (TensorRT, PyTo… | 98 | 10939 | stable |
| cumulo-autumn/StreamDiffusion StreamDiffusion is a Python pipeline for real-time interactive diffusion-based image generation, achieving 100+ fps on modern GPUs. It opti… | 17 | 10806 | active |
| OpenVINO OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su… | 95 | 10740 | stable |
| facebookresearch/xformers xFormers is a PyTorch-based library of hackable, optimized Transformer building blocks with custom CUDA kernels for fast, memory-efficient … | 88 | 10542 | active |
| skypilot-org/skypilot SkyPilot is an open-source AI compute platform that unifies fragmented infrastructure (Kubernetes, Slurm, VMs, 20+ clouds) into a single po… | 97 | 10529 | active |
| deepseek-ai/DeepEP DeepEP is a high-performance GPU communication library for expert parallelism (EP) in MoE training and inference, providing high-throughput… | 57 | 10066 | active |
| OpenRLHF/OpenRLHF OpenRLHF is a high-performance, production-ready open-source RLHF framework built on Ray + vLLM + DeepSpeed for scalable reinforcement lear… | 89 | 9956 | active |
| bentoml/BentoML BentoML is a Python framework for building online serving systems for AI apps and model inference, turning model inference scripts into RES… | 94 | 8808 | stable |
| bitsandbytes-foundation/bitsandbytes bitsandbytes is a Python library providing k-bit quantization primitives for PyTorch, enabling 8-bit (LLM.int8()) and 4-bit (QLoRA) quantiz… | 99 | 8439 | active |
| FlashML-org/FreeToken FreeToken is an edge-native Mixture-of-Experts (MoE) LLM serving engine that runs frontier-scale open-weight models on consumer hardware by… | 68 | 8370 | active |
| THUDM/slime slime is an open-source LLM post-training framework for reinforcement learning scaling, connecting Megatron-based training with SGLang-base… | 80 | 8261 | active |
| isaac-sim/IsaacLab Isaac Lab is a GPU-accelerated open-source framework for robot learning built on NVIDIA Isaac Sim, unifying workflows like reinforcement le… | 92 | 7966 | active |
| EricLBuehler/mistral.rs mistral.rs is a fast and flexible LLM inference engine written in Rust, supporting many model families with quantization (ISQ/UQFF, GGUF), … | 92 | 7628 | active |
| EleutherAI/gpt-neox GPT-NeoX is EleutherAI's library for training large-scale autoregressive transformer language models on GPUs, built on NVIDIA's Megatron an… | 62 | 7459 | active |
| naver/dust3r DUSt3R is the official PyTorch implementation of a CVPR 2024 model that performs dense, unconstrained stereo and multi-view 3D reconstructi… | 45 | 7288 | active |
| PaddlePaddle/Paddle-Lite Paddle Lite is a high-performance, lightweight deep learning inference engine from Baidu's PaddlePaddle ecosystem, designed for mobile, emb… | 58 | 7273 | active |
| kohya-ss/sd-scripts A collection of Python training, generation, and utility scripts for Stable Diffusion and other image generation models, most widely used f… | 89 | 7210 | active |
| BVLC/caffe Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp… | 23 | 34556 | maintenance |
| TheTom/turboquant_plus TurboQuant+ is a Python reference implementation of the TurboQuant KV cache compression method (ICLR 2026), using PolarQuant codebooks and … | 56 | 7006 | active |
| iperov/DeepFaceLive DeepFaceLive is a real-time face-swap application for PC streaming and video calls, using trained face models (DFM) applied to webcam or vi… | 10 | 31011 | maintenance |
| rtqichen/torchdiffeq torchdiffeq is a PyTorch library of differentiable ordinary differential equation (ODE) solvers, best known as the canonical implementation… | 39 | 6477 | stable |
| tensorflow/serving TensorFlow Serving is a flexible, high-performance serving system for machine learning models designed for production environments. It mana… | 86 | 6360 | stable |
| PufferAI/PufferLib PufferLib is a fast, open-source reinforcement learning library that trains tiny, super-human models in seconds, achieving 1M+ environment … | 81 | 6306 | active |
| google-ai-edge/LiteRT-LM LiteRT-LM is Google's production-ready, high-performance open-source framework for running large language models on edge devices, built as … | 86 | 6298 | active |
| ByteDance-Seed/Depth-Anything-3 Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,… | 59 | 6213 | active |
| harfbuzz/harfbuzz HarfBuzz is a text shaping engine that converts Unicode text into properly positioned glyph output for any writing system, supporting OpenT… | 99 | 6039 | stable |
| volcano-sh/volcano Volcano is a CNCF-hosted, Kubernetes-native batch scheduling system that extends kube-scheduler for high-performance workloads like AI/ML t… | 98 | 5899 | stable |
| Michael-A-Kuykendall/shimmy Shimmy is a single-binary, OpenAI-compatible LLM inference server written in pure Rust, running GGUF models on a WebGPU-based engine (Airfr… | 84 | 5808 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| nuclio/nuclio Nuclio is a high-performance open-source serverless (FaaS) platform for real-time event and data processing, deployable standalone via Dock… | 95 | 5750 | stable |
| pjreddie/darknet Darknet is an open-source neural network framework written in C and CUDA, best known as the original home of the YOLO real-time object dete… | 32 | 26492 | maintenance |
| NVIDIA/DALI NVIDIA DALI is a GPU-accelerated data loading and preprocessing library with optimized building blocks and an execution engine for deep lea… | 92 | 5734 | active |
| areal-project/AReaL AReaL is a large-scale asynchronous reinforcement learning system that bridges foundation model training with agent-based applications, sup… | 87 | 5696 | active |
| pytorch/torchtitan torchtitan is a PyTorch-native platform for large-scale training of generative AI models, offering a clean-room implementation of PyTorch's… | 79 | 5667 | active |
| fla-org/flash-linear-attention A PyTorch library providing hardware-efficient implementations of emerging sequence model architectures, including linear attention, sparse… | 88 | 5627 | active |
| nerfstudio-project/gsplat gsplat is an open-source Python library with CUDA-accelerated, differentiable rasterization of Gaussians, based on 3D Gaussian Splatting fo… | 70 | 5589 | active |
| gpustack/gpustack GPUStack is an open-source GPU cluster manager for AI model serving that orchestrates inference engines like vLLM, SGLang, and TensorRT-LLM… | 90 | 5560 | active |
| newton-physics/newton Newton is an open-source, GPU-accelerated physics simulation engine built on NVIDIA Warp, targeting roboticists and simulation researchers.… | 86 | 5539 | active |
| mosaicml/composer Composer is an open-source PyTorch-based deep learning training library by MosaicML (now Databricks) for training neural networks faster an… | 65 | 5495 | active |
| LaurentMazare/tch-rs tch-rs is a Rust crate providing thin bindings to the C++ API of PyTorch (libtorch), staying close to the original API. It enables tensor o… | 67 | 5479 | active |
| beclab/Olares Olares is an open-source personal cloud operating system built on Kubernetes that turns your own hardware into a self-hosted AI platform fo… | 89 | 5242 | active |
| transformerlab/transformerlab-app Transformer Lab is an open-source desktop application (built with Electron and Python) that provides a unified GUI for training, fine-tunin… | 84 | 5179 | active |
| h2oai/h2o-llmstudio H2O LLM Studio is a framework and no-code GUI for fine-tuning state-of-the-art large language models, built by H2O.ai. It supports LoRA and… | 96 | 5172 | active |
| hiyouga/EasyR1 EasyR1 is an efficient, scalable reinforcement learning training framework for large language models and vision-language models, built as a… | 65 | 5129 | active |
| arrayfire/arrayfire ArrayFire is a general-purpose tensor/numerical computing library for C, C++, and Python that accelerates array operations on GPUs (CUDA, O… | 57 | 4902 | stable |
| buxuku/SmartSub SmartSub (妙幕) is a free, open-source cross-platform desktop application that provides an end-to-end subtitle and dubbing pipeline: speech-t… | 89 | 4778 | active |
| mindspore-ai/mindspore MindSpore is an open-source deep learning framework for training and inference across mobile, edge, and cloud scenarios. It provides automa… | 32 | 4700 | active |
| RLinf/RLinf RLinf is an open-source, flexible and scalable reinforcement learning training infrastructure for embodied AI (vision-language-action model… | 74 | 4655 | active |
| Tencent/TNN TNN is a high-performance, lightweight deep learning inference framework developed by Tencent Youtu Lab, supporting mobile, desktop, and se… | 32 | 4648 | active |
| OAID/Tengine Tengine is a lightweight, high-performance, modular deep learning inference engine developed by OPEN AI LAB for embedded and edge devices. … | 27 | 4531 | active |
| NVlabs/tiny-cuda-nn A small, self-contained C++/CUDA framework for training and querying neural networks, featuring a lightning-fast fully fused MLP and a vers… | 58 | 4528 | active |
| Project-HAMi/HAMi HAMi (Heterogeneous AI Computing Virtualization Middleware) is a CNCF Incubating, Kubernetes-native GPU virtualization and scheduling middl… | 95 | 4438 | active |
| iperov/DeepFaceLab DeepFaceLab is the leading open-source Windows application for creating deepfakes, allowing users to swap, de-age, or replace faces and hea… | 10 | 19292 | maintenance |
| ModelTC/LightLLM LightLLM is a Python-based LLM inference and serving framework designed for lightweight deployment, easy scalability, and high throughput. … | 84 | 4243 | active |
| hao-ai-lab/FastVideo FastVideo is a unified Python framework for post-training and real-time inference of video diffusion models, covering data preprocessing, f… | 80 | 4076 | active |
| FedML-AI/FedML FedML (TensorOpera) is a unified Python library for large-scale distributed training, model serving, and federated learning across GPU clou… | 45 | 4062 | active |
| zml/zml ZML is a production LLM inference stack written in Zig, built on MLIR and OpenXLA, that compiles models to run at peak performance across N… | 72 | 4003 | active |
| isaac-sim/IsaacSim NVIDIA Isaac Sim is an open-source robotics simulation application built on NVIDIA Omniverse for developing, simulating, and testing AI-dri… | 74 | 3960 | active |
| Nunchaku Nunchaku is a high-performance inference engine for 4-bit quantized diffusion models (and LLMs) based on the SVDQuant technique from an ICL… | 65 | 3937 | active |
| iree-org/iree IREE is an MLIR-based end-to-end machine learning compiler and runtime that lowers models from frameworks like PyTorch, TensorFlow, JAX, an… | 87 | 3901 | active |
| kaldi-asr/kaldi Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea… | 52 | 15469 | maintenance |
| thu-ml/TurboDiffusion TurboDiffusion is a Python framework that accelerates end-to-end video diffusion model generation by 100-200x using SageAttention, Sparse-L… | 59 | 3623 | active |
| MrNeRF/LichtFeld-Studio LichtFeld Studio is a native open-source desktop application for 3D Gaussian Splatting that combines training, real-time inspection, splat … | 92 | 3594 | active |
| ashvardanian/StringZilla StringZilla is a high-performance string processing library for C, C++, Python, Rust, Swift, JS, and Go that uses SIMD, SWAR, and GPU instr… | 94 | 3542 | active |
| NVIDIA/TransformerEngine Transformer Engine is an NVIDIA library for accelerating Transformer model training and inference on NVIDIA GPUs using low-precision format… | 99 | 3504 | active |
| NVIDIA/Model-Optimizer NVIDIA Model Optimizer (ModelOpt) is a Python library of state-of-the-art model optimization techniques including quantization, pruning, di… | 91 | 3488 | active |
| guandeh17/Self-Forcing Official implementation of Self Forcing, a training method for autoregressive video diffusion models that simulates inference during traini… | 37 | 3488 | active |
| huggingface/optimum Optimum is a Hugging Face library that extends Transformers, Diffusers, timm, and Sentence Transformers with hardware-specific optimization… | 98 | 3469 | active |
| alibaba/ROLL ROLL is an open-source reinforcement learning library from Alibaba for training large language models at scale, supporting algorithms like … | 78 | 3374 | active |
| xororz/local-dream A free, open-source Android app for running Stable Diffusion locally with Snapdragon NPU acceleration, also supporting CPU/GPU inference. I… | 82 | 3346 | active |
| Mesh-LLM/mesh-llm Mesh LLM is a Rust-based distributed LLM inference runtime that pools GPUs and memory across machines into a single OpenAI-compatible API. … | 77 | 3305 | active |
| Jittor/jittor Jittor is a high-performance deep learning framework from Tsinghua University based on just-in-time (JIT) compilation and meta-operators, w… | 67 | 3229 | active |
| NVIDIA/physicsnemo NVIDIA PhysicsNeMo is an open-source Python deep-learning framework for building, training, fine-tuning, and inferring physics AI models us… | 89 | 3198 | active |
| LeelaChessZero/lc0 Lc0 is an open-source, UCI-compliant chess engine that plays chess using neural networks trained via AlphaZero-style self-play reinforcemen… | 65 | 3193 | active |
| ARM-software/ComputeLibrary Arm's Compute Library is a C++ collection of over 100 low-level machine learning and computer vision functions optimized for Arm Cortex-A/N… | 96 | 3183 | active |
| pytorch/TensorRT Torch-TensorRT is a compiler library that accelerates PyTorch model inference on NVIDIA GPUs using TensorRT. It supports just-in-time compi… | 94 | 2986 | active |