Ross ROSS = Recommend OSS · open-source software intelligence for agents

openai/prm800k resource

800,000 step-level correctness labels on LLM solutions to MATH problems observed · 2026-08-28

github.com/openai/prm800k · Python · MIT (permissive) · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 0
  • Release rhythm 35
  • Longevity 88

Flags: no_releases archived

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1238
  • days_rel: n/a
  • days_push: 1189
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2150 stars · 129 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

PRM800K is OpenAI's process supervision dataset containing 800,000 step-level correctness labels for model-generated solutions to MATH dataset problems. It accompanies the paper 'Let's Verify Step by Step' and includes raw labels plus labeler instructions.

Use cases

  • train a process reward model for math reasoning
  • get step-level correctness labels for LLM solutions
  • research process supervision versus outcome supervision
  • fine-tune LLMs on human-labeled reasoning steps
  • evaluate LLM mathematical problem solving
  • study human feedback data collection for RLHF

When to choose

  • you need step-level human feedback data for training reward models
  • you are researching process supervision for mathematical reasoning
  • you want a benchmark dataset for LLM solution correctness

When to avoid

  • you need a ready-to-use reward model rather than raw labels
  • you need labels for domains other than math problems
  • you want actively maintained tooling rather than a static dataset

Facets

dataset · maturity maintenance

machine-learning llm-training data-science large-language-models machine-learning mathematics python cross-platform process-supervision reward-models mathematical-reasoning step-level-labels llm-evaluation human-annotation natural-language-processing

1 source

Member repositories

RepositoryRoleHealth v2
openai/prm800kmain10

For agents

markdown · JSON · MCP: product_card(name="openai/prm800k")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem