function: gpu-computing
559 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| pytorch/ao TorchAO is a PyTorch-native library for model optimization through quantization and sparsity. It supports quantizing weights, gradients, op… | 89 | 2957 | active |
| luminal-ai/luminal Luminal is a high-performance general-purpose ML inference compiler written in Rust that lowers models to a minimal 15-op dataflow IR and c… | 78 | 2956 | active |
| bghira/SimpleTuner SimpleTuner is a Python fine-tuning toolkit for image, video, and audio diffusion models built on Hugging Face Diffusers. It provides a web… | 92 | 2912 | active |
| mitsuba-renderer/mitsuba3 Mitsuba 3 is a research-oriented, retargetable rendering system for forward and inverse light transport simulation, written in C++17 on top… | 94 | 2899 | active |
| mujocolab/mjlab mjlab is a Python framework that combines Isaac Lab's manager-based API with MuJoCo Warp, a GPU-accelerated version of MuJoCo, for reinforc… | 85 | 2837 | active |
| superlinked/sie SIE (Superlinked Inference Engine) is an open-source, self-hosted inference server and production cluster that serves 100+ open models (emb… | 92 | 2830 | active |
| pytorch/xla PyTorch/XLA is a Python package that connects the PyTorch deep learning framework to XLA devices such as Google Cloud TPUs via the XLA deep… | 69 | 2803 | active |
| NVlabs/stylegan2 The official TensorFlow implementation of StyleGAN2, NVIDIA's improved style-based generative adversarial network for high-quality uncondit… | 32 | 11184 | maintenance |
| ModelTC/LightX2V LightX2V is a lightweight, high-performance inference framework for image and video generation, supporting tasks like text-to-video, image-… | 64 | 2733 | active |
| huggingface/text-generation-inference Text Generation Inference (TGI) is a Rust, Python and gRPC toolkit for deploying and serving large language models with high performance, p… | 10 | 10889 | maintenance |
| NVlabs/LongLive LongLive is an NVIDIA research framework providing parallel training and inference infrastructure for real-time long video generation, usin… | 60 | 2563 | active |
| jolibrain/deepdetect DeepDetect is an open-source deep learning runtime, CLI, and REST server written in C++ for training and inference across images, text, tab… | 95 | 2551 | active |
| yifan123/flow_grpo Flow-GRPO is the official PyTorch implementation of a NeurIPS 2025 paper that trains flow matching models (e.g., SD3.5, FLUX.1, Qwen-Image,… | 56 | 2498 | active |
| antirez/h3.c A native C inference engine for the MiniMax H3 model on Apple Silicon, using Metal for GPU acceleration. It generates video (and audio) fro… | 56 | 2443 | active |
| google/tunix Tunix is a lightweight JAX-based library for post-training large language models, supporting supervised fine-tuning, preference optimizatio… | 83 | 2415 | active |
| microsoft/Olive Olive is Microsoft's AI model optimization toolkit for the ONNX Runtime, automating finetuning, conversion, quantization, and compression o… | 91 | 2382 | active |
| openlake-project/openlake OpenLake is a high-performance distributed storage engine written in Rust (built on io_uring, RDMA, and GPUDirect) designed to feed GPUs du… | 81 | 2330 | active |
| radixark/miles Miles is an open-source, enterprise-grade reinforcement learning framework for large-scale LLM and VLM post-training, forked from and co-ev… | 72 | 2263 | active |
| mfem/mfem MFEM is a lightweight, modular C++ library for finite element discretization of PDEs, supporting arbitrary high-order element spaces, adapt… | 74 | 2226 | stable |
| dstackai/dstack dstack is an open-source, vendor-agnostic control plane for GPU provisioning and orchestration that works across GPU clouds, Kubernetes, an… | 95 | 2221 | active |
| learnhouse/learnhouse LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com… | 99 | 2203 | active |
| NovaSky-AI/SkyRL SkyRL is a modular full-stack reinforcement learning library for post-training large language models, combining a training framework (skyrl… | 80 | 2201 | active |
| MetalPetal/MetalPetal MetalPetal is a GPU-accelerated image and video processing framework built on Apple's Metal API. It provides an image/filter/render pipelin… | 23 | 2179 | active |
| google-deepmind/mujoco_playground MuJoCo Playground is an open-source Python library of GPU-accelerated robot learning environments built on MuJoCo MJX and MuJoCo Warp. It s… | 76 | 2168 | active |
| nv-tlabs/vipe ViPE is an open-source video processing engine from NVIDIA that estimates camera intrinsics, camera motion, and dense near-metric depth map… | 80 | 2092 | active |
| patrick-kidger/diffrax Diffrax is a JAX-based library providing numerical differential equation solvers for ODEs, SDEs, and CDEs. It is fully autodifferentiable a… | 80 | 2089 | active |
| marcoslucianops/DeepStream-Yolo A collection of configuration files, parsers, and conversion utilities for running YOLO-family object detection models on NVIDIA DeepStream… | 61 | 2054 | active |
| NUS-HPC-AI-Lab/VideoSys VideoSys is an open-source Python library providing easy and efficient infrastructure for video generation, supporting training, inference,… | 45 | 2022 | active |
| chapel-lang/chapel Chapel is a modern open-source programming language designed for productive parallel computing at scale, with first-class support for task … | 87 | 2017 | active |
| tdrussell/diffusion-pipe A Python training script for fine-tuning diffusion models (image and video generation) using DeepSpeed pipeline parallelism across multiple… | 67 | 2015 | active |
| mil-tokyo/webdnn WebDNN is a framework for running deep neural network inference directly in the web browser, accepting ONNX models without Python preproces… | 65 | 1999 | active |
| not-an-aardvark/lucky-commit A Rust CLI tool that amends git commits with whitespace until the commit hash starts with a desired prefix (default '0000000'). It uses GPU… | 41 | 1985 | active |
| PrimeIntellect-ai/prime-rl prime-rl is a Python framework for large-scale, fully asynchronous reinforcement learning training of language models, built on FSDP2 for t… | 87 | 1975 | active |
| NVIDIA-NeMo/RL NeMo RL is NVIDIA's open-source post-training library for scaling reinforcement learning methods (GRPO, PPO, DPO, SFT, distillation) on LLM… | 81 | 1961 | active |
| AbdBarho/stable-diffusion-webui-docker A Docker Compose setup that runs Stable Diffusion locally with popular web UIs like AUTOMATIC1111, ComfyUI, and InvokeAI. It packages model… | 23 | 7309 | maintenance |
| flexflow/flexflow-train FlexFlow Train is a deep learning framework that accelerates distributed DNN training by automatically searching for efficient parallelizat… | 67 | 1898 | active |
| laugh12321/TensorRT-YOLO A C++/Python deployment toolkit for running YOLO-family models (YOLOv3 through YOLO26) on NVIDIA GPUs using TensorRT, with custom plugins, … | 63 | 1880 | active |
| nndeploy/nndeploy nndeploy is an easy-to-use, high-performance AI deployment framework written in C++ with Python bindings. It provides a visual drag-and-dro… | 87 | 1868 | active |
| NVIDIA-AI-IOT/Lidar_AI_Solution NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil… | 72 | 1867 | active |
| NVlabs/stylegan3 Official PyTorch implementation of StyleGAN3 (Alias-Free GANs), a state-of-the-art generative adversarial network for high-fidelity image s… | 32 | 6943 | maintenance |
| Tavris1/ComfyUI-Easy-Install ComfyUI-Easy-Install is a portable one-click installer for ComfyUI that bundles an EZi Desktop application, requiring no manual Python or G… | 85 | 1831 | active |
| nazdridoy/kokoro-tts A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w… | 88 | 1816 | active |
| triple-mu/YOLOv8-TensorRT A library for running YOLOv8 inference accelerated with NVIDIA TensorRT, supporting detection, segmentation, pose estimation, oriented boun… | 75 | 1804 | active |
| koide3/glim GLIM is a versatile and extensible point cloud-based 3D localization and mapping (SLAM) framework written in C++. It performs direct multi-… | 76 | 1758 | active |
| beam-cloud/beta9 Beam (beta9) is an open-source serverless runtime for AI workloads, providing GPU inference endpoints, isolated sandboxes for running untru… | 89 | 1755 | active |
| NVIDIA/FasterTransformer NVIDIA's highly optimized C++/CUDA library for fast inference of Transformer-based models such as BERT, GPT, and encoder-decoder models, wi… | 23 | 6447 | maintenance |
| BindsNET/bindsnet BindsNET is a Python package for simulating spiking neural networks (SNNs) built on PyTorch tensor functionality, running on CPUs or GPUs. … | 84 | 1695 | active |
| tkarras/progressive_growing_of_gans Official TensorFlow implementation of the ICLR 2018 NVIDIA paper 'Progressive Growing of GANs', which trains generators and discriminators … | 32 | 6179 | maintenance |
| 2swap/swaptube SwapTube is a C++ framework for programmatically rendering YouTube videos, built on FFMPEG with custom graphics code above the encoding lay… | 74 | 1652 | active |
| tenstorrent/tt-metal TT-Metal is Tenstorrent's open-source software stack containing TT-NN, a Python and C++ neural network operator library, and TT-Metalium, a… | 97 | 1640 | active |
| gomlx/gomlx GoMLX is an accelerated machine learning and math framework for Go, comparable to PyTorch/JAX/TensorFlow. It offers differentiable operator… | 95 | 1621 | active |
| gorgonia/gorgonia Gorgonia is a Go library for machine learning that lets you define and evaluate mathematical equations over multidimensional arrays using a… | 23 | 5929 | maintenance |
| alibaba/Pai-Megatron-Patch Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and V… | 56 | 1591 | active |
| Tencent-Hunyuan/HunyuanWorld-Voyager HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ… | 52 | 1590 | active |
| Tencent-Hunyuan/HY-WorldPlay HY-WorldPlay (HY-World 1.5) is Tencent Hunyuan's open-source framework for interactive 3D world modeling, generating explorable 3D scenes f… | 55 | 1586 | active |
| pytorch/FBGEMM FBGEMM is a collection of highly optimized low-precision matrix multiplication and convolution kernels for server-side deep learning infere… | 93 | 1584 | active |
| NVlabs/sionna Sionna is an open-source, GPU-accelerated, differentiable Python library from NVIDIA for research on communication systems. It comprises Si… | 83 | 1577 | active |
| Xilinx/brevitas Brevitas is a PyTorch library for neural network quantization supporting both post-training quantization (PTQ) and quantization-aware train… | 91 | 1567 | active |
| amandaghassaei/gpu-io gpu-io is a TypeScript WebGL library for composing GPU-accelerated computing workflows in the browser. It handles WebGL state management, s… | 32 | 1482 | active |
| Soul-AILab/SoulX-FlashTalk SoulX-FlashTalk is a 14B audio-driven talking avatar model that streams infinite real-time video from a reference image and audio, achievin… | 58 | 1479 | active |
| Lightning-AI/lightning-thunder Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a … | 72 | 1469 | active |
| FeiYull/TensorRT-Alpha A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl… | 32 | 1460 | active |
| tensorflow/tpu A collection of reference models and tools for training machine learning models on Google Cloud TPUs, maintained as a public mirror by the … | 72 | 5278 | maintenance |
| deepseek-ai/EPLB EPLB is DeepSeek's open-source Expert Parallelism Load Balancer for Mixture-of-Experts models. It computes balanced expert replication and … | 26 | 1424 | active |
| trilinos/Trilinos Trilinos is a collection of C++ libraries and an object-oriented software framework for solving large-scale, complex multi-physics engineer… | 95 | 1418 | stable |
| mratsim/Arraymancer Arraymancer is a fast, ergonomic N-dimensional tensor (ndarray) library written in Nim, inspired by NumPy and PyTorch. It provides CPU, CUD… | 61 | 1407 | active |
| davideberly/GeometricTools The Geometric Tools Engine (GTE) is a C++14 collection of source code for computing in mathematics, geometry, graphics, image analysis, and… | 76 | 1382 | active |
| OpenPPL/ppl.nn PPLNN is a high-performance deep-learning inference engine written in C++ that runs ONNX models on x86 CPUs and NVIDIA GPUs, with a dedicat… | 32 | 1367 | active |
| NVIDIA-AI-IOT/torch2trt torch2trt is a Python library that converts PyTorch models to TensorRT engines using the TensorRT Python API, with a simple single-function… | 23 | 4878 | maintenance |
| k2-fsa/k2 k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)… | 64 | 1352 | active |
| amandaghassaei/OrigamiSimulator A realtime WebGL web application that simulates how any origami crease pattern folds, solving all creases simultaneously via GPU fragment s… | 56 | 1339 | active |
| jonathan-laurent/AlphaZero.jl A generic, simple, and fast Julia implementation of DeepMind's AlphaZero algorithm for training game-playing agents via self-play and MCTS.… | 64 | 1333 | active |
| facebookincubator/AITemplate AITemplate is a Python framework that compiles deep neural networks into high-performance CUDA (NVIDIA) or HIP (AMD) C++ code for fast fp16… | 66 | 4724 | maintenance |
| nomadkaraoke/python-audio-separator A Python package and CLI that separates audio files into stems (vocals, instrumental, drums, bass, etc.) using pre-trained models from Ulti… | 89 | 1327 | active |
| ACEsuit/mace MACE is a Python library implementing fast and accurate machine learning interatomic potentials using higher-order equivariant message pass… | 89 | 1324 | active |
| meta-pytorch/segment-anything-fast A fast, batched offline inference-oriented fork of Meta's Segment Anything (SAM) image segmentation model. It applies optimizations like bf… | 45 | 1321 | active |
| anvaka/fieldplay Field Play is a browser-based WebGL application for exploring and visualizing vector fields by animating thousands of GPU-driven particles.… | 67 | 1318 | active |
| derrian-distro/LoRA_Easy_Training_Scripts A PySide6 desktop GUI that wraps Kohya's sd-scripts to simplify training LoRA, LoCon, and other LoRA-type models for Stable Diffusion. It s… | 33 | 1307 | active |
| plaidml/plaidml PlaidML is a portable tensor compiler that enables deep learning on hardware (especially GPUs and embedded devices) not well supported by m… | 10 | 4566 | maintenance |
| BytedTsinghua-SIA/CUDA-Agent CUDA-Agent is a large-scale agentic reinforcement learning system from ByteDance Seed and Tsinghua that trains LLMs to generate high-perfor… | 56 | 1256 | active |
| Neroued/ninfer NInfer is a from-scratch C++/CUDA inference engine optimized for maximum single-GPU performance on a narrow, explicitly registered set of Q… | 58 | 1252 | active |
| ModelCloud/GPTQModel GPTQModel is a production-ready Python toolkit for quantizing (compressing) large language models using GPTQ, AWQ, and related methods, wit… | 91 | 1248 | active |
| Ksuriuri/index-tts-vllm A reimplementation of IndexTTS's GPT model inference using vLLM, providing significantly faster text-to-speech generation with a web UI and… | 58 | 1235 | active |
| chengzeyi/Comfy-WaveSpeed A ComfyUI custom node plugin that acts as an all-in-one inference optimization solution for diffusion models, built around First Block Cach… | 66 | 1231 | stable |
| fastgs/FastGS FastGS is a general acceleration framework for 3D Gaussian Splatting that trains scenes in roughly 100 seconds using multi-view consistent … | 49 | 1221 | active |
| antvis/G G is a flexible 2D/3D rendering engine for visualization, serving as the underlying graphics engine of the AntV charting ecosystem. It adap… | 73 | 1211 | active |
| dendenxu/fast-gaussian-rasterization A drop-in replacement for diff-gaussian-rasterization that renders 3D Gaussian Splatting scenes using a geometry-shader-based GPU pipeline … | 19 | 1202 | active |
| steelbrain/ffmpeg-over-ip A client/server tool that lets applications use GPU-accelerated ffmpeg on a remote machine over a single TCP connection, without GPU passth… | 83 | 1176 | active |
| Soul-AILab/SoulX-LiveAct SoulX-LiveAct is the official inference code for a real-time human animation framework that generates lifelike, audio/multimodal-controlled… | 54 | 1176 | active |
| Linaom1214/TensorRT-For-YOLO-Series A Python and C++ toolkit for running YOLO-series object detection models (YOLOv3 through YOLOv12, YOLOX) with NVIDIA TensorRT, including ON… | 44 | 1162 | active |
| open-gigaai/giga-world-1 GigaWorld-1 is an open-source framework providing training, inference, data processing, checkpoint conversion, and LoRA merge workflows for… | 54 | 1147 | active |
| NVIDIA-Omniverse/kit-app-template A template toolkit from NVIDIA for building GPU-accelerated, OpenUSD-based 3D applications with the Omniverse Kit SDK. It provides pre-conf… | 75 | 1090 | active |
| flagos-ai/FlagGems FlagGems is a high-performance operator library for large language models written in the Triton language, providing backend-neutral GPU ker… | 89 | 1089 | active |
| mlc-ai/web-stable-diffusion A project that compiles and runs Stable Diffusion text-to-image models entirely inside web browsers using WebGPU and WebAssembly, with no s… | 30 | 3721 | maintenance |
| NVlabs/Fast-dLLM NVIDIA's official implementation of Fast-dLLM, a family of training-free and fine-tuning-based acceleration techniques for diffusion-based … | 57 | 1082 | active |
| open-gigaai/giga-train GigaTrain is an efficient and scalable Python training framework for large AI models, supporting distributed multi-GPU/multi-node execution… | 62 | 1072 | active |
| bytedance/SandboxFusion A secure, self-hosted code sandbox service from ByteDance that runs and judges code generated by LLMs across 20+ programming languages via … | 63 | 1060 | active |
| XYZ-AI-Lab/axrl AxisRL is an agentic reinforcement learning post-training framework for large language models, built on SGLang for high-throughput rollout … | 55 | 1056 | active |
| arcee-ai/DistillKit DistillKit is an open-source Python toolkit for knowledge distillation of large language models, supporting both online and offline distill… | 60 | 1047 | active |
| LuisaGroup/LuisaCompute LuisaCompute is a high-performance cross-platform computing framework for graphics and beyond, featuring a C++-embedded DSL for GPU kernel … | 77 | 1043 | active |