Ross ROSS = Recommend OSS · open-source software intelligence for agents

kubeflow/trainer

Distributed AI Model Training and LLM Fine-Tuning on Kubernetes observed · 2026-08-28

github.com/kubeflow/trainer · homepage · Go · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

95/100

  • Activity 99
  • Release rhythm 86
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 62
  • age_days: 3353
  • days_rel: 15
  • days_push: 7
  • n_releases_24m: 12

Full methodology

Adoption not part of the score

2198 stars · 1034 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Kubeflow Trainer is a Kubernetes-native platform for distributed AI model training and LLM fine-tuning across frameworks like PyTorch, JAX, TensorFlow, HuggingFace, and XGBoost. It orchestrates multi-node, multi-GPU jobs (including MPI workloads) via TrainJob custom resources with a Python SDK.

Use cases

  • fine-tune LLMs on Kubernetes with multiple GPUs
  • run distributed PyTorch training jobs across nodes
  • train models with JAX or TensorFlow on a cluster
  • orchestrate MPI workloads on Kubernetes
  • scale HuggingFace model training with DeepSpeed
  • manage ML training jobs as Kubernetes resources

When to choose

  • you already run Kubernetes and need scalable distributed training
  • you want LLM fine-tuning with multi-node, multi-GPU orchestration
  • you need framework-agnostic training runtimes (PyTorch, JAX, XGBoost, MPI)

When to avoid

  • you train small models on a single machine without a cluster
  • you have no Kubernetes infrastructure or ops expertise
  • you only need inference/serving rather than training

Facets

framework · maturity active

machine-learning llm-training gpu-computing workflow-automation deployment machine-learning deep-learning large-language-models microservices python go cloud distributed-training llm-fine-tuning pytorch jax tensorflow huggingface mpi mlops trainjob kueue kubernetes docker

1 source

Member repositories

RepositoryRoleHealth v2
kubeflow/trainermain95

For agents

markdown · JSON · MCP: product_card(name="kubeflow/trainer")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem