Ross ROSS = Recommend OSS · open-source software intelligence for agents

alibaba/Pai-Megatron-Patch

The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud. observed · 2026-08-28

github.com/alibaba/Pai-Megatron-Patch · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

56/100

  • Activity 57
  • Release rhythm 42
  • Longevity 78
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 31
  • age_days: 1094
  • days_rel: 306
  • days_push: 261
  • n_releases_24m: 14

Full methodology

Adoption not part of the score

1591 stars · 233 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Pai-Megatron-Patch is an open-source deep learning training toolkit from Alibaba Cloud for large-scale training and inference of LLMs and VLMs built on the Megatron framework. It provides optimized training recipes, examples, and integrations (Megatron-Core, ChatLearn, verl) for models like Qwen, DeepSeek, and Moonlight.

Use cases

  • fine-tune large language models with Megatron efficiently
  • train vision-language models like Qwen2.5-VL at scale
  • run RLHF or GRPO post-training on models like DeepSeek-R1
  • pretrain models larger than 10 billion parameters with high GPU utilization
  • convert HuggingFace checkpoints to Megatron format and back
  • train Qwen3 models on Alibaba Cloud PAI

When to choose

  • you need maximum training efficiency for 10B+ parameter models
  • you want ready-made recipes for Qwen, DeepSeek, or Moonlight architectures
  • you are training or fine-tuning LLMs/VLMs on Alibaba Cloud PAI or multi-GPU clusters
  • you need RLHF/GRPO training pipelines via ChatLearn or verl integration

When to avoid

  • you only need to fine-tune small models where HuggingFace Transformers or LoRA suffices
  • you want a simple single-GPU training script
  • you are not working with NVIDIA GPU clusters or Megatron-style parallelism
  • you need a managed no-code training platform rather than a Python toolkit

Facets

library · maturity active

llm-training deep-learning machine-learning gpu-computing large-language-models deep-learning machine-learning gpu-computing python cloud megatron llm-fine-tuning vlm-training qwen deepseek distributed-training reinforcement-learning alibaba-cloud pai gpu linux docker

1 source

Member repositories

RepositoryRoleHealth v2
alibaba/Pai-Megatron-Patchmain56

For agents

markdown · JSON · MCP: product_card(name="alibaba/Pai-Megatron-Patch")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem