# NVIDIA/deepops

Tools for building GPU clusters

Repository: https://github.com/NVIDIA/deepops
Canonical: https://ross.abutalabs.com/products/nvidia-deepops
Language: Shell
License: BSD-3-Clause
License Family: permissive
Last push: 2026-08-24T18:12:53+00:00

## Health v2 (maintenance only)
Score: 93/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 81, longevity 100
- inputs: {"age_days": 3023, "days_push": 9, "days_rel": 50, "gap_med": 49, "n_releases_24m": 2}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1468, forks 355 (observed 2026-08-28T04:04:48.946287+00:00)

## What it is
NVIDIA DeepOps is a collection of Ansible playbooks and scripts for deploying GPU clusters running Kubernetes or Slurm, including drivers, Docker, and the NVIDIA Container Runtime. It encapsulates deployment best practices for NVIDIA DGX systems and modular GPU infrastructure setups.

## Use cases
- set up an on-prem GPU cluster of DGX servers
- deploy Kubernetes with GPU support on bare metal
- install Slurm as a batch scheduler on GPU nodes
- deploy Kubeflow on an existing Kubernetes cluster
- configure NVIDIA drivers and container runtime on a single GPU machine
- attach NFS storage to a GPU cluster
- build a hybrid Slurm and Kubernetes cluster

## When to choose
- you are standing up a new GPU cluster with Kubernetes or Slurm
- you manage NVIDIA DGX systems and want validated end-to-end deployment tooling
- you need GPU drivers, Docker, and NVIDIA Container Runtime installed automatically
- you want Ansible-based, modular infrastructure automation for HPC or ML clusters

## When to avoid
- you only need a managed cloud Kubernetes service with GPUs
- your nodes run unsupported or legacy operating systems you cannot validate
- you need a GUI-driven cluster manager rather than scripts and playbooks
- you are deploying on non-NVIDIA accelerators

## Facets
- artifact type: infra-config
- maturity: active
- function: infrastructure-as-code, container-orchestration, deployment, configuration-management, gpu-computing
- domain: infrastructure-as-code, gpu-computing, cloud-computing, microservices
- platform: self-hosted
- tags: gpu-clusters, slurm, kubernetes, ansible, nvidia-dgx, hpc, cluster-deployment, nvidia, devops, containers, linux, docker, gpu

## Member repositories
- NVIDIA/deepops (main) score 93

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:48.946287+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:34:55.402978+00:00, confidence not recorded.
  - readme: https://github.com/NVIDIA/deepops (fetched 2026-08-28T04:04:48.946287+00:00, sha 697e9c7b7eff)
- Data as of 2026-08-30T08:39:29.467469+00:00.
