Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: gpu-computing

559 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
pytorch/ao
TorchAO is a PyTorch-native library for model optimization through quantization and sparsity. It supports quantizing weights, gradients, op…
892957active
luminal-ai/luminal
Luminal is a high-performance general-purpose ML inference compiler written in Rust that lowers models to a minimal 15-op dataflow IR and c…
782956active
bghira/SimpleTuner
SimpleTuner is a Python fine-tuning toolkit for image, video, and audio diffusion models built on Hugging Face Diffusers. It provides a web…
922912active
mitsuba-renderer/mitsuba3
Mitsuba 3 is a research-oriented, retargetable rendering system for forward and inverse light transport simulation, written in C++17 on top…
942899active
mujocolab/mjlab
mjlab is a Python framework that combines Isaac Lab's manager-based API with MuJoCo Warp, a GPU-accelerated version of MuJoCo, for reinforc…
852837active
superlinked/sie
SIE (Superlinked Inference Engine) is an open-source, self-hosted inference server and production cluster that serves 100+ open models (emb…
922830active
pytorch/xla
PyTorch/XLA is a Python package that connects the PyTorch deep learning framework to XLA devices such as Google Cloud TPUs via the XLA deep…
692803active
NVlabs/stylegan2
The official TensorFlow implementation of StyleGAN2, NVIDIA's improved style-based generative adversarial network for high-quality uncondit…
3211184maintenance
ModelTC/LightX2V
LightX2V is a lightweight, high-performance inference framework for image and video generation, supporting tasks like text-to-video, image-…
642733active
huggingface/text-generation-inference
Text Generation Inference (TGI) is a Rust, Python and gRPC toolkit for deploying and serving large language models with high performance, p…
1010889maintenance
NVlabs/LongLive
LongLive is an NVIDIA research framework providing parallel training and inference infrastructure for real-time long video generation, usin…
602563active
jolibrain/deepdetect
DeepDetect is an open-source deep learning runtime, CLI, and REST server written in C++ for training and inference across images, text, tab…
952551active
yifan123/flow_grpo
Flow-GRPO is the official PyTorch implementation of a NeurIPS 2025 paper that trains flow matching models (e.g., SD3.5, FLUX.1, Qwen-Image,…
562498active
antirez/h3.c
A native C inference engine for the MiniMax H3 model on Apple Silicon, using Metal for GPU acceleration. It generates video (and audio) fro…
562443active
google/tunix
Tunix is a lightweight JAX-based library for post-training large language models, supporting supervised fine-tuning, preference optimizatio…
832415active
microsoft/Olive
Olive is Microsoft's AI model optimization toolkit for the ONNX Runtime, automating finetuning, conversion, quantization, and compression o…
912382active
openlake-project/openlake
OpenLake is a high-performance distributed storage engine written in Rust (built on io_uring, RDMA, and GPUDirect) designed to feed GPUs du…
812330active
radixark/miles
Miles is an open-source, enterprise-grade reinforcement learning framework for large-scale LLM and VLM post-training, forked from and co-ev…
722263active
mfem/mfem
MFEM is a lightweight, modular C++ library for finite element discretization of PDEs, supporting arbitrary high-order element spaces, adapt…
742226stable
dstackai/dstack
dstack is an open-source, vendor-agnostic control plane for GPU provisioning and orchestration that works across GPU clouds, Kubernetes, an…
952221active
learnhouse/learnhouse
LearnHouse is a next-generation open-source learning management system (LMS) for creating, sharing, and selling educational content. It com…
992203active
NovaSky-AI/SkyRL
SkyRL is a modular full-stack reinforcement learning library for post-training large language models, combining a training framework (skyrl…
802201active
MetalPetal/MetalPetal
MetalPetal is a GPU-accelerated image and video processing framework built on Apple's Metal API. It provides an image/filter/render pipelin…
232179active
google-deepmind/mujoco_playground
MuJoCo Playground is an open-source Python library of GPU-accelerated robot learning environments built on MuJoCo MJX and MuJoCo Warp. It s…
762168active
nv-tlabs/vipe
ViPE is an open-source video processing engine from NVIDIA that estimates camera intrinsics, camera motion, and dense near-metric depth map…
802092active
patrick-kidger/diffrax
Diffrax is a JAX-based library providing numerical differential equation solvers for ODEs, SDEs, and CDEs. It is fully autodifferentiable a…
802089active
marcoslucianops/DeepStream-Yolo
A collection of configuration files, parsers, and conversion utilities for running YOLO-family object detection models on NVIDIA DeepStream…
612054active
NUS-HPC-AI-Lab/VideoSys
VideoSys is an open-source Python library providing easy and efficient infrastructure for video generation, supporting training, inference,…
452022active
chapel-lang/chapel
Chapel is a modern open-source programming language designed for productive parallel computing at scale, with first-class support for task …
872017active
tdrussell/diffusion-pipe
A Python training script for fine-tuning diffusion models (image and video generation) using DeepSpeed pipeline parallelism across multiple…
672015active
mil-tokyo/webdnn
WebDNN is a framework for running deep neural network inference directly in the web browser, accepting ONNX models without Python preproces…
651999active
not-an-aardvark/lucky-commit
A Rust CLI tool that amends git commits with whitespace until the commit hash starts with a desired prefix (default '0000000'). It uses GPU…
411985active
PrimeIntellect-ai/prime-rl
prime-rl is a Python framework for large-scale, fully asynchronous reinforcement learning training of language models, built on FSDP2 for t…
871975active
NVIDIA-NeMo/RL
NeMo RL is NVIDIA's open-source post-training library for scaling reinforcement learning methods (GRPO, PPO, DPO, SFT, distillation) on LLM…
811961active
AbdBarho/stable-diffusion-webui-docker
A Docker Compose setup that runs Stable Diffusion locally with popular web UIs like AUTOMATIC1111, ComfyUI, and InvokeAI. It packages model…
237309maintenance
flexflow/flexflow-train
FlexFlow Train is a deep learning framework that accelerates distributed DNN training by automatically searching for efficient parallelizat…
671898active
laugh12321/TensorRT-YOLO
A C++/Python deployment toolkit for running YOLO-family models (YOLOv3 through YOLO26) on NVIDIA GPUs using TensorRT, with custom plugins, …
631880active
nndeploy/nndeploy
nndeploy is an easy-to-use, high-performance AI deployment framework written in C++ with Python bindings. It provides a visual drag-and-dro…
871868active
NVIDIA-AI-IOT/Lidar_AI_Solution
NVIDIA's collection of GPU-accelerated Lidar AI inference solutions for autonomous driving, including optimized implementations of PointPil…
721867active
NVlabs/stylegan3
Official PyTorch implementation of StyleGAN3 (Alias-Free GANs), a state-of-the-art generative adversarial network for high-fidelity image s…
326943maintenance
Tavris1/ComfyUI-Easy-Install
ComfyUI-Easy-Install is a portable one-click installer for ComfyUI that bundles an EZi Desktop application, requiring no manual Python or G…
851831active
nazdridoy/kokoro-tts
A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w…
881816active
triple-mu/YOLOv8-TensorRT
A library for running YOLOv8 inference accelerated with NVIDIA TensorRT, supporting detection, segmentation, pose estimation, oriented boun…
751804active
koide3/glim
GLIM is a versatile and extensible point cloud-based 3D localization and mapping (SLAM) framework written in C++. It performs direct multi-…
761758active
beam-cloud/beta9
Beam (beta9) is an open-source serverless runtime for AI workloads, providing GPU inference endpoints, isolated sandboxes for running untru…
891755active
NVIDIA/FasterTransformer
NVIDIA's highly optimized C++/CUDA library for fast inference of Transformer-based models such as BERT, GPT, and encoder-decoder models, wi…
236447maintenance
BindsNET/bindsnet
BindsNET is a Python package for simulating spiking neural networks (SNNs) built on PyTorch tensor functionality, running on CPUs or GPUs. …
841695active
tkarras/progressive_growing_of_gans
Official TensorFlow implementation of the ICLR 2018 NVIDIA paper 'Progressive Growing of GANs', which trains generators and discriminators …
326179maintenance
2swap/swaptube
SwapTube is a C++ framework for programmatically rendering YouTube videos, built on FFMPEG with custom graphics code above the encoding lay…
741652active
tenstorrent/tt-metal
TT-Metal is Tenstorrent's open-source software stack containing TT-NN, a Python and C++ neural network operator library, and TT-Metalium, a…
971640active
gomlx/gomlx
GoMLX is an accelerated machine learning and math framework for Go, comparable to PyTorch/JAX/TensorFlow. It offers differentiable operator…
951621active
gorgonia/gorgonia
Gorgonia is a Go library for machine learning that lets you define and evaluate mathematical equations over multidimensional arrays using a…
235929maintenance
alibaba/Pai-Megatron-Patch
Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and V…
561591active
Tencent-Hunyuan/HunyuanWorld-Voyager
HunyuanWorld-Voyager is a video diffusion framework from Tencent Hunyuan that generates world-consistent RGBD video and 3D point-cloud sequ…
521590active
Tencent-Hunyuan/HY-WorldPlay
HY-WorldPlay (HY-World 1.5) is Tencent Hunyuan's open-source framework for interactive 3D world modeling, generating explorable 3D scenes f…
551586active
pytorch/FBGEMM
FBGEMM is a collection of highly optimized low-precision matrix multiplication and convolution kernels for server-side deep learning infere…
931584active
NVlabs/sionna
Sionna is an open-source, GPU-accelerated, differentiable Python library from NVIDIA for research on communication systems. It comprises Si…
831577active
Xilinx/brevitas
Brevitas is a PyTorch library for neural network quantization supporting both post-training quantization (PTQ) and quantization-aware train…
911567active
amandaghassaei/gpu-io
gpu-io is a TypeScript WebGL library for composing GPU-accelerated computing workflows in the browser. It handles WebGL state management, s…
321482active
Soul-AILab/SoulX-FlashTalk
SoulX-FlashTalk is a 14B audio-driven talking avatar model that streams infinite real-time video from a reference image and audio, achievin…
581479active
Lightning-AI/lightning-thunder
Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a …
721469active
FeiYull/TensorRT-Alpha
A C++/CUDA library providing TensorRT-accelerated deployment for 30+ popular computer vision models including YOLOv3-v8, YOLOv8-Pose/Seg/Cl…
321460active
tensorflow/tpu
A collection of reference models and tools for training machine learning models on Google Cloud TPUs, maintained as a public mirror by the …
725278maintenance
deepseek-ai/EPLB
EPLB is DeepSeek's open-source Expert Parallelism Load Balancer for Mixture-of-Experts models. It computes balanced expert replication and …
261424active
trilinos/Trilinos
Trilinos is a collection of C++ libraries and an object-oriented software framework for solving large-scale, complex multi-physics engineer…
951418stable
mratsim/Arraymancer
Arraymancer is a fast, ergonomic N-dimensional tensor (ndarray) library written in Nim, inspired by NumPy and PyTorch. It provides CPU, CUD…
611407active
davideberly/GeometricTools
The Geometric Tools Engine (GTE) is a C++14 collection of source code for computing in mathematics, geometry, graphics, image analysis, and…
761382active
OpenPPL/ppl.nn
PPLNN is a high-performance deep-learning inference engine written in C++ that runs ONNX models on x86 CPUs and NVIDIA GPUs, with a dedicat…
321367active
NVIDIA-AI-IOT/torch2trt
torch2trt is a Python library that converts PyTorch models to TensorRT engines using the TensorRT Python API, with a simple single-function…
234878maintenance
k2-fsa/k2
k2 is a C++/CUDA library with Python bindings that implements differentiable Finite State Automaton (FSA) and Finite State Transducer (FST)…
641352active
amandaghassaei/OrigamiSimulator
A realtime WebGL web application that simulates how any origami crease pattern folds, solving all creases simultaneously via GPU fragment s…
561339active
jonathan-laurent/AlphaZero.jl
A generic, simple, and fast Julia implementation of DeepMind's AlphaZero algorithm for training game-playing agents via self-play and MCTS.…
641333active
facebookincubator/AITemplate
AITemplate is a Python framework that compiles deep neural networks into high-performance CUDA (NVIDIA) or HIP (AMD) C++ code for fast fp16…
664724maintenance
nomadkaraoke/python-audio-separator
A Python package and CLI that separates audio files into stems (vocals, instrumental, drums, bass, etc.) using pre-trained models from Ulti…
891327active
ACEsuit/mace
MACE is a Python library implementing fast and accurate machine learning interatomic potentials using higher-order equivariant message pass…
891324active
meta-pytorch/segment-anything-fast
A fast, batched offline inference-oriented fork of Meta's Segment Anything (SAM) image segmentation model. It applies optimizations like bf…
451321active
anvaka/fieldplay
Field Play is a browser-based WebGL application for exploring and visualizing vector fields by animating thousands of GPU-driven particles.…
671318active
derrian-distro/LoRA_Easy_Training_Scripts
A PySide6 desktop GUI that wraps Kohya's sd-scripts to simplify training LoRA, LoCon, and other LoRA-type models for Stable Diffusion. It s…
331307active
plaidml/plaidml
PlaidML is a portable tensor compiler that enables deep learning on hardware (especially GPUs and embedded devices) not well supported by m…
104566maintenance
BytedTsinghua-SIA/CUDA-Agent
CUDA-Agent is a large-scale agentic reinforcement learning system from ByteDance Seed and Tsinghua that trains LLMs to generate high-perfor…
561256active
Neroued/ninfer
NInfer is a from-scratch C++/CUDA inference engine optimized for maximum single-GPU performance on a narrow, explicitly registered set of Q…
581252active
ModelCloud/GPTQModel
GPTQModel is a production-ready Python toolkit for quantizing (compressing) large language models using GPTQ, AWQ, and related methods, wit…
911248active
Ksuriuri/index-tts-vllm
A reimplementation of IndexTTS's GPT model inference using vLLM, providing significantly faster text-to-speech generation with a web UI and…
581235active
chengzeyi/Comfy-WaveSpeed
A ComfyUI custom node plugin that acts as an all-in-one inference optimization solution for diffusion models, built around First Block Cach…
661231stable
fastgs/FastGS
FastGS is a general acceleration framework for 3D Gaussian Splatting that trains scenes in roughly 100 seconds using multi-view consistent …
491221active
antvis/G
G is a flexible 2D/3D rendering engine for visualization, serving as the underlying graphics engine of the AntV charting ecosystem. It adap…
731211active
dendenxu/fast-gaussian-rasterization
A drop-in replacement for diff-gaussian-rasterization that renders 3D Gaussian Splatting scenes using a geometry-shader-based GPU pipeline …
191202active
steelbrain/ffmpeg-over-ip
A client/server tool that lets applications use GPU-accelerated ffmpeg on a remote machine over a single TCP connection, without GPU passth…
831176active
Soul-AILab/SoulX-LiveAct
SoulX-LiveAct is the official inference code for a real-time human animation framework that generates lifelike, audio/multimodal-controlled…
541176active
Linaom1214/TensorRT-For-YOLO-Series
A Python and C++ toolkit for running YOLO-series object detection models (YOLOv3 through YOLOv12, YOLOX) with NVIDIA TensorRT, including ON…
441162active
open-gigaai/giga-world-1
GigaWorld-1 is an open-source framework providing training, inference, data processing, checkpoint conversion, and LoRA merge workflows for…
541147active
NVIDIA-Omniverse/kit-app-template
A template toolkit from NVIDIA for building GPU-accelerated, OpenUSD-based 3D applications with the Omniverse Kit SDK. It provides pre-conf…
751090active
flagos-ai/FlagGems
FlagGems is a high-performance operator library for large language models written in the Triton language, providing backend-neutral GPU ker…
891089active
mlc-ai/web-stable-diffusion
A project that compiles and runs Stable Diffusion text-to-image models entirely inside web browsers using WebGPU and WebAssembly, with no s…
303721maintenance
NVlabs/Fast-dLLM
NVIDIA's official implementation of Fast-dLLM, a family of training-free and fine-tuning-based acceleration techniques for diffusion-based …
571082active
open-gigaai/giga-train
GigaTrain is an efficient and scalable Python training framework for large AI models, supporting distributed multi-GPU/multi-node execution…
621072active
bytedance/SandboxFusion
A secure, self-hosted code sandbox service from ByteDance that runs and judges code generated by LLMs across 20+ programming languages via …
631060active
XYZ-AI-Lab/axrl
AxisRL is an agentic reinforcement learning post-training framework for large language models, built on SGLang for high-throughput rollout …
551056active
arcee-ai/DistillKit
DistillKit is an open-source Python toolkit for knowledge distillation of large language models, supporting both online and offline distill…
601047active
LuisaGroup/LuisaCompute
LuisaCompute is a high-performance cross-platform computing framework for graphics and beyond, featuring a C++-embedded DSL for GPU kernel …
771043active

← prev page 5 / 6 next →