Ross ROSS = Recommend OSS · open-source software intelligence for agents

cfregly/ai-performance-engineering resource

Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning. observed · 2026-08-28

github.com/cfregly/ai-performance-engineering · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

63/100

  • Activity 98
  • Release rhythm 35
  • Longevity 35

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 499
  • days_rel: n/a
  • days_push: 12
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1863 stars · 260 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Code, labs, and resources accompanying the O'Reilly book 'AI Systems Performance Engineering', covering GPU optimization, distributed training, and inference scaling. It includes thousands of lines of PyTorch and CUDA C++ examples plus profiling case studies for modern NVIDIA GPUs.

Use cases

  • learn how to profile and optimize GPU workloads with Nsight and PyTorch profilers
  • understand distributed training parallelism strategies like FSDP, TP, PP, and MoE
  • optimize high-throughput LLM inference with vLLM, SGLang, and TensorRT-LLM
  • write custom GPU kernels with Triton and the PyTorch compiler stack
  • reduce cost per token when serving large models
  • hands-on labs for AI systems performance tuning

When to choose

  • you are an ML or systems engineer building or operating training/inference at scale
  • you want empirical, profiling-driven GPU optimization techniques
  • you need to learn disaggregated prefill/decode serving and KV-cache management
  • you want a structured book plus runnable code examples

When to avoid

  • you need a production-ready tool or framework rather than educational material
  • you work exclusively on non-NVIDIA hardware
  • you are a beginner looking for basic ML tutorials rather than performance engineering

Facets

learning-resource · maturity active

benchmarking gpu-computing llm-inference llm-training machine-learning deep-learning gpu-computing machine-learning deep-learning large-language-models performance developer-tools tutorials python cross-platform gpu-optimization distributed-training inference-serving profiling cuda pytorch triton vllm tensorrt-llm oreilly-book labs gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
cfregly/ai-performance-engineeringmain63

For agents

markdown · JSON · MCP: product_card(name="cfregly/ai-performance-engineering")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem