# walkinglabs/hands-on-modern-rl

🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.

Repository: https://github.com/walkinglabs/hands-on-modern-rl
Canonical: https://ross.abutalabs.com/products/hands-on-modern-rl
Homepage: https://walkinglabs.github.io/hands-on-modern-rl/
Language: Python
License: NOASSERTION
License Family: other
Topics: agent, agentic, agentic-ai, agentic-rl, dpo, grpo, llm, llm-alignment, pytorch, rlhf, tutorial, ppo, reinforcement, reinforcement-learning, rl, sft
Last push: 2026-08-25T14:10:43+00:00

## Health v2 (maintenance only)
Score: 78/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 89, longevity 10
- inputs: {"age_days": 146, "days_push": 8, "days_rel": 76, "gap_med": 0.0, "n_releases_24m": 7}
- flags: young, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 4124, forks 297 (observed 2026-08-28T04:08:35.950104+00:00)

## What it is
An open-source hands-on curriculum/book covering modern reinforcement learning, from MDPs and policy optimization (PPO) to LLM alignment (RLHF, DPO, GRPO, RLVR) and agentic systems. It includes experiment code, notebooks, and online environments, published as a VitePress site and PDF under CC BY-NC-SA 4.0.

## Use cases
- learn reinforcement learning from scratch with hands-on code
- understand how RLHF and DPO align LLMs
- study GRPO and RLVR for reasoning models
- find a tutorial bridging classic RL to agentic AI
- run CartPole and classic RL experiments in notebooks
- learn policy optimization with PyTorch

## When to choose
- you want a structured, code-driven path from basic RL to LLM alignment and agents
- you prefer runnable notebooks and online experiments alongside theory
- you need coverage of modern post-training methods like GRPO, DPO, and RLVR

## When to avoid
- you need a production RL training framework rather than a learning resource
- you require rigorously peer-reviewed content - the course is AI-assisted and not fully reviewed
- you need a permissively licensed codebase - content is CC BY-NC-SA (non-commercial)

## Facets
- artifact type: learning-resource
- maturity: active
- function: reinforcement-learning, llm-training, agent-framework, machine-learning
- domain: reinforcement-learning, large-language-models, tutorials, education
- platform: python, cross-platform
- tags: rlhf, grpo, dpo, ppo, rlvr, llm-alignment, curriculum, pytorch, open-course, ai-agents

## Member repositories
- walkinglabs/hands-on-modern-rl (main) score 78

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:35.950104+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:23:05.023194+00:00, confidence not recorded.
  - readme: https://github.com/walkinglabs/hands-on-modern-rl (fetched 2026-08-28T04:08:35.950104+00:00, sha 089201b634e1)
  - homepage: https://walkinglabs.github.io/hands-on-modern-rl/ (fetched 2026-08-29T09:14:16.875037+00:00, sha 6e106a726d62)
- Data as of 2026-08-30T08:39:29.467469+00:00.
