NVIDIA/deepops resource
Tools for building GPU clusters observed · 2026-08-28
Health v2 · maintenance only
93/100
- Activity 99
- Release rhythm 81
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 49
- age_days: 3023
- days_rel: 50
- days_push: 9
- n_releases_24m: 2
Adoption not part of the score
1468 stars · 355 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
NVIDIA DeepOps is a collection of Ansible playbooks and scripts for deploying GPU clusters running Kubernetes or Slurm, including drivers, Docker, and the NVIDIA Container Runtime. It encapsulates deployment best practices for NVIDIA DGX systems and modular GPU infrastructure setups.
Use cases
- set up an on-prem GPU cluster of DGX servers
- deploy Kubernetes with GPU support on bare metal
- install Slurm as a batch scheduler on GPU nodes
- deploy Kubeflow on an existing Kubernetes cluster
- configure NVIDIA drivers and container runtime on a single GPU machine
- attach NFS storage to a GPU cluster
- build a hybrid Slurm and Kubernetes cluster
When to choose
- you are standing up a new GPU cluster with Kubernetes or Slurm
- you manage NVIDIA DGX systems and want validated end-to-end deployment tooling
- you need GPU drivers, Docker, and NVIDIA Container Runtime installed automatically
- you want Ansible-based, modular infrastructure automation for HPC or ML clusters
When to avoid
- you only need a managed cloud Kubernetes service with GPUs
- your nodes run unsupported or legacy operating systems you cannot validate
- you need a GUI-driven cluster manager rather than scripts and playbooks
- you are deploying on non-NVIDIA accelerators
Facets
infra-config · maturity active
infrastructure-as-code container-orchestration deployment configuration-management gpu-computing infrastructure-as-code gpu-computing cloud-computing microservices self-hosted gpu-clusters slurm kubernetes ansible nvidia-dgx hpc cluster-deployment nvidia devops containers linux docker gpu
1 source
- readme: https://github.com/NVIDIA/deepops · fetched 2026-08-28 · 697e9c7b7eff
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| NVIDIA/deepops | main | 93 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem