Ross ROSS = Recommend OSS · open-source software intelligence for agents

RUC-NLPIR/ARPO

[ICLR 2026] Agentic Reinforced Policy Optimization (ARPO) observed · 2026-08-28

github.com/RUC-NLPIR/ARPO · Python observed · 2026-08-28

Health v2 · maintenance only

54/100

  • Activity 98
  • Release rhythm 13
  • Longevity 29

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 407
  • days_rel: 367
  • days_push: 13
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

1109 stars · 60 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

ARPO (Agentic Reinforced Policy Optimization) is a reinforcement learning algorithm and training framework for LLM agents, published at ICLR 2026. It provides training code, models, and datasets for improving multi-turn agentic reasoning and tool use via RL.

Use cases

  • train llm agents with reinforcement learning
  • improve multi-turn tool use of a language model
  • reproduce ARPO paper results
  • fine-tune Qwen or Llama models for agentic search
  • apply RL post-training to deep search agents

When to choose

  • you need a research-grade agentic RL training algorithm for LLMs
  • you want to RL-train models for multi-turn tool calling or deep search
  • you want to build on ICLR 2026 ARPO/AEPO methods

When to avoid

  • you need a production-supported training framework with commercial support
  • you need simple SFT-only fine-tuning without RL
  • you lack GPU infrastructure for large-model RL training

Facets

library · maturity active

llm-training reinforcement-learning agent-framework large-language-models reinforcement-learning machine-learning python rlhf policy-optimization agentic-rl iclr-2026 research-code llm-post-training ai-agents gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
RUC-NLPIR/ARPOmain54

For agents

markdown · JSON · MCP: product_card(name="RUC-NLPIR/ARPO")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem