# alibaba/Pai-Megatron-Patch

The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.

Repository: https://github.com/alibaba/Pai-Megatron-Patch
Canonical: https://ross.abutalabs.com/products/pai-megatron-patch
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2025-12-15T14:13:03+00:00

## Health v2 (maintenance only)
Score: 56/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 57, release rhythm 42, longevity 78
- inputs: {"age_days": 1094, "days_push": 261, "days_rel": 306, "gap_med": 31, "n_releases_24m": 14}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1591, forks 233 (observed 2026-08-28T04:05:08.655212+00:00)

## What it is
Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and VLMs built on the Megatron framework. It provides optimized training recipes, examples, and integrations (Megatron-Core, ChatLearn, verl) for models like Qwen, DeepSeek, and Moonlight.

## Use cases
- fine-tune large language models with Megatron efficiently
- train vision-language models like Qwen2.5-VL at scale
- run RLHF or GRPO post-training on models like DeepSeek-R1
- pretrain models larger than 10 billion parameters with high GPU utilization
- convert HuggingFace checkpoints to Megatron format and back
- train Qwen3 models on Alibaba Cloud PAI

## When to choose
- you need maximum training efficiency for 10B+ parameter models
- you want ready-made recipes for Qwen, DeepSeek, or Moonlight architectures
- you are training or fine-tuning LLMs/VLMs on Alibaba Cloud PAI or multi-GPU clusters
- you need RLHF/GRPO training pipelines via ChatLearn or verl integration

## When to avoid
- you only need to fine-tune small models where HuggingFace Transformers or LoRA suffices
- you want a simple single-GPU training script
- you are not working with NVIDIA GPU clusters or Megatron-style parallelism
- you need a managed no-code training platform rather than a Python toolkit

## Facets
- artifact type: library
- maturity: active
- function: llm-training, deep-learning, machine-learning, gpu-computing
- domain: large-language-models, deep-learning, machine-learning, gpu-computing
- platform: python, cloud
- tags: megatron, llm-fine-tuning, vlm-training, qwen, deepseek, distributed-training, reinforcement-learning, alibaba-cloud, pai, gpu, linux, docker

## Member repositories
- alibaba/Pai-Megatron-Patch (main) score 56

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:08.655212+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:54:43.392828+00:00, confidence not recorded.
  - readme: https://github.com/alibaba/Pai-Megatron-Patch (fetched 2026-08-28T04:05:08.655212+00:00, sha 9da0bb60de41)
- Data as of 2026-08-30T08:39:29.467469+00:00.
