kubeflow/trainer
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes observed · 2026-08-28
Health v2 · maintenance only
95/100
- Activity 99
- Release rhythm 86
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 62
- age_days: 3353
- days_rel: 15
- days_push: 7
- n_releases_24m: 12
Adoption not part of the score
2198 stars · 1034 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Kubeflow Trainer is a Kubernetes-native platform for distributed AI model training and LLM fine-tuning across frameworks like PyTorch, JAX, TensorFlow, HuggingFace, and XGBoost. It orchestrates multi-node, multi-GPU jobs (including MPI workloads) via TrainJob custom resources with a Python SDK.
Use cases
- fine-tune LLMs on Kubernetes with multiple GPUs
- run distributed PyTorch training jobs across nodes
- train models with JAX or TensorFlow on a cluster
- orchestrate MPI workloads on Kubernetes
- scale HuggingFace model training with DeepSpeed
- manage ML training jobs as Kubernetes resources
When to choose
- you already run Kubernetes and need scalable distributed training
- you want LLM fine-tuning with multi-node, multi-GPU orchestration
- you need framework-agnostic training runtimes (PyTorch, JAX, XGBoost, MPI)
When to avoid
- you train small models on a single machine without a cluster
- you have no Kubernetes infrastructure or ops expertise
- you only need inference/serving rather than training
Facets
framework · maturity active
machine-learning llm-training gpu-computing workflow-automation deployment machine-learning deep-learning large-language-models microservices python go cloud distributed-training llm-fine-tuning pytorch jax tensorflow huggingface mpi mlops trainjob kueue kubernetes docker
1 source
- readme: https://github.com/kubeflow/trainer · fetched 2026-08-28 · f58edb23ae12
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| kubeflow/trainer | main | 95 |
For agents
markdown · JSON · MCP: product_card(name="kubeflow/trainer")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem