Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: gpu-computing

559 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
abiosoft/colima
Colima is a CLI tool that provides container runtimes (Docker, Containerd, Incus) on macOS and Linux with minimal setup, built on Lima. It …
9530526active
modular/modular
Modular Platform hosting the MAX AI serving framework and the Mojo systems programming language. It provides an OpenAI-compatible inference…
9229225active
MLX
MLX is an array computation framework for machine learning on Apple silicon, developed by Apple ML research. It offers NumPy-like Python AP…
9428172active
Stability-AI/generative-models
Stability AI's official repository of generative diffusion models, including Stable Diffusion, Stable Video, and SV4D 2.0 for image, video,…
4527271active
ApolloAuto/apollo
Apollo is an open-source autonomous driving platform providing a high-performance, modular software stack for developing, testing, and depl…
5726807active
PaddlePaddle/Paddle
PaddlePaddle is an industrial-grade deep learning framework written in C++ with Python APIs, supporting high-performance single-machine and…
8424062active
Tencent/ncnn
ncnn is a high-performance neural network inference framework written in C++ and optimized for mobile, embedded, and desktop deployment. It…
8623753stable
verl-project/verl
verl (Volcano Engine Reinforcement Learning) is a flexible, production-ready RL post-training library for large language models, open-sourc…
8423145active
microsoft/onnxruntime
ONNX Runtime is a cross-platform, high-performance machine-learning accelerator for running inference and training on ONNX models. It suppo…
9921654stable
k4yt3x/video2x
Video2X is a machine learning-based video super-resolution and frame interpolation framework written in C/C++. It upscales videos and image…
6621263active
huggingface/candle
Candle is a minimalist machine learning framework for Rust focused on performance and ease of use, with CPU and CUDA GPU support. It ships …
7320955active
NVIDIA/Megatron-LM
NVIDIA's GPU-optimized library for training large transformer models at scale, comprising Megatron-LM (reference training scripts) and Mega…
9917615active
alibaba/MNN
MNN is a lightweight, high-performance deep learning inference engine developed by Alibaba, supporting LLMs, vision, audio, and multimodal …
9315973active
tracel-ai/burn
Burn is a Rust-based tensor library and deep learning framework supporting training and inference through a unified API. It JIT-compiles te…
8915816active
GeeeekExplorer/nano-vllm
A lightweight vLLM-style LLM inference engine implemented from scratch in about 1,200 lines of Python. It offers fast offline inference wit…
5415164active
facebookresearch/vggt
VGGT (Visual Geometry Grounded Transformer) is a feed-forward transformer model from Meta AI and Oxford VGG that infers 3D geometry—camera …
5814292active
Eclipse Deeplearning4J
Eclipse Deeplearning4J is an open-source deep learning framework and ecosystem for the JVM, including the ND4J linear algebra library, the …
7714246active
Open3D
Open3D is an open-source C++ and Python library for 3D data processing, offering data structures, algorithms, and pipelines for point cloud…
6713913active
apache/tvm
Apache TVM is an open machine learning compiler framework that takes pre-trained models and compiles them into optimized, deployable module…
9013691active
openwall/john
John the Ripper jumbo is an open-source offline password cracker supporting hundreds of hash and cipher types, from Unix and Windows passwo…
7513543active
NVIDIA/TensorRT
NVIDIA TensorRT is an SDK for high-performance deep learning inference on NVIDIA GPUs, comprising an inference compiler that converts train…
9413293stable
modelscope/DiffSynth-Studio
DiffSynth-Studio is an open-source diffusion model engine from the ModelScope community that integrates mainstream image, video, and audio …
7913003active
cupy/cupy
CuPy is a NumPy/SciPy-compatible array library for GPU-accelerated computing with Python, running on NVIDIA CUDA or AMD ROCm. It acts as a …
9512278stable
LMCache/LMCache
LMCache is a KV cache management layer for LLM inference that stores, compresses, and reuses KV caches across requests, sessions, and servi…
8611460active
triton-inference-server/server
NVIDIA Triton Inference Server is an open-source inference serving software that deploys AI models from multiple frameworks (TensorRT, PyTo…
9810939stable
cumulo-autumn/StreamDiffusion
StreamDiffusion is a Python pipeline for real-time interactive diffusion-based image generation, achieving 100+ fps on modern GPUs. It opti…
1710806active
OpenVINO
OpenVINO is an open-source toolkit from Intel for optimizing and deploying deep learning inference across CPU, GPU, and NPU hardware. It su…
9510740stable
facebookresearch/xformers
xFormers is a PyTorch-based library of hackable, optimized Transformer building blocks with custom CUDA kernels for fast, memory-efficient …
8810542active
skypilot-org/skypilot
SkyPilot is an open-source AI compute platform that unifies fragmented infrastructure (Kubernetes, Slurm, VMs, 20+ clouds) into a single po…
9710529active
deepseek-ai/DeepEP
DeepEP is a high-performance GPU communication library for expert parallelism (EP) in MoE training and inference, providing high-throughput…
5710066active
OpenRLHF/OpenRLHF
OpenRLHF is a high-performance, production-ready open-source RLHF framework built on Ray + vLLM + DeepSpeed for scalable reinforcement lear…
899956active
bentoml/BentoML
BentoML is a Python framework for building online serving systems for AI apps and model inference, turning model inference scripts into RES…
948808stable
bitsandbytes-foundation/bitsandbytes
bitsandbytes is a Python library providing k-bit quantization primitives for PyTorch, enabling 8-bit (LLM.int8()) and 4-bit (QLoRA) quantiz…
998439active
FlashML-org/FreeToken
FreeToken is an edge-native Mixture-of-Experts (MoE) LLM serving engine that runs frontier-scale open-weight models on consumer hardware by…
688370active
THUDM/slime
slime is an open-source LLM post-training framework for reinforcement learning scaling, connecting Megatron-based training with SGLang-base…
808261active
isaac-sim/IsaacLab
Isaac Lab is a GPU-accelerated open-source framework for robot learning built on NVIDIA Isaac Sim, unifying workflows like reinforcement le…
927966active
EricLBuehler/mistral.rs
mistral.rs is a fast and flexible LLM inference engine written in Rust, supporting many model families with quantization (ISQ/UQFF, GGUF), …
927628active
EleutherAI/gpt-neox
GPT-NeoX is EleutherAI's library for training large-scale autoregressive transformer language models on GPUs, built on NVIDIA's Megatron an…
627459active
naver/dust3r
DUSt3R is the official PyTorch implementation of a CVPR 2024 model that performs dense, unconstrained stereo and multi-view 3D reconstructi…
457288active
PaddlePaddle/Paddle-Lite
Paddle Lite is a high-performance, lightweight deep learning inference engine from Baidu's PaddlePaddle ecosystem, designed for mobile, emb…
587273active
kohya-ss/sd-scripts
A collection of Python training, generation, and utility scripts for Stable Diffusion and other image generation models, most widely used f…
897210active
BVLC/caffe
Caffe is a fast, modular deep learning framework developed by Berkeley AI Research, focused on convolutional neural networks for vision, sp…
2334556maintenance
TheTom/turboquant_plus
TurboQuant+ is a Python reference implementation of the TurboQuant KV cache compression method (ICLR 2026), using PolarQuant codebooks and …
567006active
iperov/DeepFaceLive
DeepFaceLive is a real-time face-swap application for PC streaming and video calls, using trained face models (DFM) applied to webcam or vi…
1031011maintenance
rtqichen/torchdiffeq
torchdiffeq is a PyTorch library of differentiable ordinary differential equation (ODE) solvers, best known as the canonical implementation…
396477stable
tensorflow/serving
TensorFlow Serving is a flexible, high-performance serving system for machine learning models designed for production environments. It mana…
866360stable
PufferAI/PufferLib
PufferLib is a fast, open-source reinforcement learning library that trains tiny, super-human models in seconds, achieving 1M+ environment …
816306active
google-ai-edge/LiteRT-LM
LiteRT-LM is Google's production-ready, high-performance open-source framework for running large language models on edge devices, built as …
866298active
ByteDance-Seed/Depth-Anything-3
Depth Anything 3 (DA3) is a transformer-based model that predicts spatially consistent depth and geometry from any number of visual inputs,…
596213active
harfbuzz/harfbuzz
HarfBuzz is a text shaping engine that converts Unicode text into properly positioned glyph output for any writing system, supporting OpenT…
996039stable
volcano-sh/volcano
Volcano is a CNCF-hosted, Kubernetes-native batch scheduling system that extends kube-scheduler for high-performance workloads like AI/ML t…
985899stable
Michael-A-Kuykendall/shimmy
Shimmy is a single-binary, OpenAI-compatible LLM inference server written in pure Rust, running GGUF models on a WebGPU-based engine (Airfr…
845808active
aidlearning/AidLearning-FrameWork
AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling…
705797active
nuclio/nuclio
Nuclio is a high-performance open-source serverless (FaaS) platform for real-time event and data processing, deployable standalone via Dock…
955750stable
pjreddie/darknet
Darknet is an open-source neural network framework written in C and CUDA, best known as the original home of the YOLO real-time object dete…
3226492maintenance
NVIDIA/DALI
NVIDIA DALI is a GPU-accelerated data loading and preprocessing library with optimized building blocks and an execution engine for deep lea…
925734active
areal-project/AReaL
AReaL is a large-scale asynchronous reinforcement learning system that bridges foundation model training with agent-based applications, sup…
875696active
pytorch/torchtitan
torchtitan is a PyTorch-native platform for large-scale training of generative AI models, offering a clean-room implementation of PyTorch's…
795667active
fla-org/flash-linear-attention
A PyTorch library providing hardware-efficient implementations of emerging sequence model architectures, including linear attention, sparse…
885627active
nerfstudio-project/gsplat
gsplat is an open-source Python library with CUDA-accelerated, differentiable rasterization of Gaussians, based on 3D Gaussian Splatting fo…
705589active
gpustack/gpustack
GPUStack is an open-source GPU cluster manager for AI model serving that orchestrates inference engines like vLLM, SGLang, and TensorRT-LLM…
905560active
newton-physics/newton
Newton is an open-source, GPU-accelerated physics simulation engine built on NVIDIA Warp, targeting roboticists and simulation researchers.…
865539active
mosaicml/composer
Composer is an open-source PyTorch-based deep learning training library by MosaicML (now Databricks) for training neural networks faster an…
655495active
LaurentMazare/tch-rs
tch-rs is a Rust crate providing thin bindings to the C++ API of PyTorch (libtorch), staying close to the original API. It enables tensor o…
675479active
beclab/Olares
Olares is an open-source personal cloud operating system built on Kubernetes that turns your own hardware into a self-hosted AI platform fo…
895242active
transformerlab/transformerlab-app
Transformer Lab is an open-source desktop application (built with Electron and Python) that provides a unified GUI for training, fine-tunin…
845179active
h2oai/h2o-llmstudio
H2O LLM Studio is a framework and no-code GUI for fine-tuning state-of-the-art large language models, built by H2O.ai. It supports LoRA and…
965172active
hiyouga/EasyR1
EasyR1 is an efficient, scalable reinforcement learning training framework for large language models and vision-language models, built as a…
655129active
arrayfire/arrayfire
ArrayFire is a general-purpose tensor/numerical computing library for C, C++, and Python that accelerates array operations on GPUs (CUDA, O…
574902stable
buxuku/SmartSub
SmartSub (妙幕) is a free, open-source cross-platform desktop application that provides an end-to-end subtitle and dubbing pipeline: speech-t…
894778active
mindspore-ai/mindspore
MindSpore is an open-source deep learning framework for training and inference across mobile, edge, and cloud scenarios. It provides automa…
324700active
RLinf/RLinf
RLinf is an open-source, flexible and scalable reinforcement learning training infrastructure for embodied AI (vision-language-action model…
744655active
Tencent/TNN
TNN is a high-performance, lightweight deep learning inference framework developed by Tencent Youtu Lab, supporting mobile, desktop, and se…
324648active
OAID/Tengine
Tengine is a lightweight, high-performance, modular deep learning inference engine developed by OPEN AI LAB for embedded and edge devices. …
274531active
NVlabs/tiny-cuda-nn
A small, self-contained C++/CUDA framework for training and querying neural networks, featuring a lightning-fast fully fused MLP and a vers…
584528active
Project-HAMi/HAMi
HAMi (Heterogeneous AI Computing Virtualization Middleware) is a CNCF Incubating, Kubernetes-native GPU virtualization and scheduling middl…
954438active
iperov/DeepFaceLab
DeepFaceLab is the leading open-source Windows application for creating deepfakes, allowing users to swap, de-age, or replace faces and hea…
1019292maintenance
ModelTC/LightLLM
LightLLM is a Python-based LLM inference and serving framework designed for lightweight deployment, easy scalability, and high throughput. …
844243active
hao-ai-lab/FastVideo
FastVideo is a unified Python framework for post-training and real-time inference of video diffusion models, covering data preprocessing, f…
804076active
FedML-AI/FedML
FedML (TensorOpera) is a unified Python library for large-scale distributed training, model serving, and federated learning across GPU clou…
454062active
zml/zml
ZML is a production LLM inference stack written in Zig, built on MLIR and OpenXLA, that compiles models to run at peak performance across N…
724003active
isaac-sim/IsaacSim
NVIDIA Isaac Sim is an open-source robotics simulation application built on NVIDIA Omniverse for developing, simulating, and testing AI-dri…
743960active
Nunchaku
Nunchaku is a high-performance inference engine for 4-bit quantized diffusion models (and LLMs) based on the SVDQuant technique from an ICL…
653937active
iree-org/iree
IREE is an MLIR-based end-to-end machine learning compiler and runtime that lowers models from frameworks like PyTorch, TensorFlow, JAX, an…
873901active
kaldi-asr/kaldi
Kaldi is a C++ toolkit for speech recognition research and development, including acoustic modeling, feature extraction, decoding, and spea…
5215469maintenance
thu-ml/TurboDiffusion
TurboDiffusion is a Python framework that accelerates end-to-end video diffusion model generation by 100-200x using SageAttention, Sparse-L…
593623active
MrNeRF/LichtFeld-Studio
LichtFeld Studio is a native open-source desktop application for 3D Gaussian Splatting that combines training, real-time inspection, splat …
923594active
ashvardanian/StringZilla
StringZilla is a high-performance string processing library for C, C++, Python, Rust, Swift, JS, and Go that uses SIMD, SWAR, and GPU instr…
943542active
NVIDIA/TransformerEngine
Transformer Engine is an NVIDIA library for accelerating Transformer model training and inference on NVIDIA GPUs using low-precision format…
993504active
NVIDIA/Model-Optimizer
NVIDIA Model Optimizer (ModelOpt) is a Python library of state-of-the-art model optimization techniques including quantization, pruning, di…
913488active
guandeh17/Self-Forcing
Official implementation of Self Forcing, a training method for autoregressive video diffusion models that simulates inference during traini…
373488active
huggingface/optimum
Optimum is a Hugging Face library that extends Transformers, Diffusers, timm, and Sentence Transformers with hardware-specific optimization…
983469active
alibaba/ROLL
ROLL is an open-source reinforcement learning library from Alibaba for training large language models at scale, supporting algorithms like …
783374active
xororz/local-dream
A free, open-source Android app for running Stable Diffusion locally with Snapdragon NPU acceleration, also supporting CPU/GPU inference. I…
823346active
Mesh-LLM/mesh-llm
Mesh LLM is a Rust-based distributed LLM inference runtime that pools GPUs and memory across machines into a single OpenAI-compatible API. …
773305active
Jittor/jittor
Jittor is a high-performance deep learning framework from Tsinghua University based on just-in-time (JIT) compilation and meta-operators, w…
673229active
NVIDIA/physicsnemo
NVIDIA PhysicsNeMo is an open-source Python deep-learning framework for building, training, fine-tuning, and inferring physics AI models us…
893198active
LeelaChessZero/lc0
Lc0 is an open-source, UCI-compliant chess engine that plays chess using neural networks trained via AlphaZero-style self-play reinforcemen…
653193active
ARM-software/ComputeLibrary
Arm's Compute Library is a C++ collection of over 100 low-level machine learning and computer vision functions optimized for Arm Cortex-A/N…
963183active
pytorch/TensorRT
Torch-TensorRT is a compiler library that accelerates PyTorch model inference on NVIDIA GPUs using TensorRT. It supports just-in-time compi…
942986active

← prev page 4 / 6 next →