Ross ROSS = Recommend OSS · open-source software intelligence for agents

google-deepmind/bsuite

bsuite is a collection of carefully-designed experiments that investigate core capabilities of a reinforcement learning (RL) agent observed · 2026-08-28

github.com/google-deepmind/bsuite · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

82/100

  • Activity 99
  • Release rhythm 51
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2588
  • days_rel: 117
  • days_push: 9
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

1555 stars · 187 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

bsuite (Behaviour Suite for Reinforcement Learning) is a collection of carefully-designed experiments from DeepMind that investigate core capabilities of reinforcement learning agents. It automates evaluation and analysis of any agent on shared benchmarks via environment-embedded logging and pre-made Jupyter notebook analysis.

Use cases

  • benchmark my RL agent on standardized capability experiments
  • evaluate reinforcement learning algorithm generalization
  • compare RL agents on exploration and memory tasks
  • analyze agent behavior with reproducible plots
  • find weaknesses in my RL algorithm's learning
  • run reproducible RL research experiments

When to choose

  • you need standardized, well-designed benchmarks to evaluate RL agent capabilities
  • you want automated logging and analysis notebooks for RL experiments
  • you are doing RL research requiring reproducible comparisons across algorithms

When to avoid

  • you need a general-purpose RL environment library like Gym for training rather than evaluation
  • you need high-throughput or large-scale environment simulation
  • you want pre-trained RL agents or algorithms rather than evaluation tooling

Facets

library · maturity stable

machine-learning reinforcement-learning benchmarking testing data-visualization reinforcement-learning machine-learning python windows reinforcement-learning-benchmark rl-evaluation deepmind agent-capabilities jupyter-notebook research-benchmark research algorithms linux macos

1 source

Member repositories

RepositoryRoleHealth v2
google-deepmind/bsuitemain82

For agents

markdown · JSON · MCP: product_card(name="google-deepmind/bsuite")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem