# microsoft/pai

Resource scheduling and cluster management for AI

Repository: https://github.com/microsoft/pai
Canonical: https://ross.abutalabs.com/products/pai
Homepage: https://openpai.readthedocs.io
Language: JavaScript
License: MIT
License Family: permissive
Topics: kubernetes, resource-management, scheduling, machine-learning, tensorflow, cluster-manager, gpu, model-training, ai, artificial-intelligence, pytorch, jupyter, chainer, cluster-management, cloud, on-premise, gpu-cluster, gpu-computing, gpu-scheduler
Last push: 2026-08-15T00:09:09+00:00

## Health v2 (maintenance only)
Score: 66/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 97, release rhythm 8, longevity 100
- inputs: {"age_days": 3264, "days_push": 19, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2688, forks 552 (observed 2026-08-28T04:07:10.824942+00:00)

## What it is
OpenPAI is an open-source AI platform from Microsoft that provides resource scheduling and cluster management for machine learning workloads on GPU clusters, built on Kubernetes. It offers a web portal, job submission SDK, and VS Code integration for training models with frameworks like TensorFlow and PyTorch.

## Use cases
- manage a shared GPU cluster for deep learning teams
- schedule TensorFlow and PyTorch training jobs on Kubernetes
- allocate GPU resources fairly across AI researchers
- run Jupyter notebooks on a dedicated AI cluster
- set up an on-premise AI training platform
- queue and monitor distributed model training jobs

## When to choose
- you need a full-featured, self-hosted GPU cluster manager for AI training
- your team shares limited GPU hardware and needs job scheduling and quotas
- you want a web portal and SDK on top of Kubernetes for ML workloads

## When to avoid
- you need active development or new features - the repo is read-only since v1.8.1 (Dec 2021)
- you prefer a modern alternative like Kubeflow, Ray, or Volcano
- you only need single-machine training without cluster scheduling

## Facets
- artifact type: service
- maturity: maintenance
- function: container-orchestration, scheduling, machine-learning, llm-training, gpu-computing, cloud, self-hosted
- domain: machine-learning, deep-learning, gpu-computing, cloud-computing, infrastructure-as-code, artificial-intelligence
- platform: self-hosted, cloud
- tags: gpu-cluster, cluster-management, resource-scheduling, model-training, jupyter, tensorflow, pytorch, on-premise, microsoft, devops, kubernetes, docker, linux

## Member repositories
- microsoft/pai (main) score 66

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:10.824942+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:17:44.217828+00:00, confidence not recorded.
  - readme: https://github.com/microsoft/pai (fetched 2026-08-28T04:07:10.824942+00:00, sha de04d3ba6d22)
- Data as of 2026-08-30T08:39:29.467469+00:00.
