# openai/following-instructions-human-feedback

Repository: https://github.com/openai/following-instructions-human-feedback
Canonical: https://ross.abutalabs.com/products/following-instructions-human-feedback
License Family: other
Archived: true
Last push: 2022-12-11T19:58:53+00:00

## Health v2 (maintenance only)
Score: 10/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 1682, "days_push": 1361, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, archived, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1259, forks 149 (observed 2026-08-28T04:04:09.716827+00:00)

## What it is
The official companion repository for OpenAI's InstructGPT paper on aligning language models with human intent via reinforcement learning from human feedback (RLHF). It contains the model card, evaluation samples, and labeling instructions rather than runnable training code.

## Use cases
- understand how RLHF aligns language models with user intent
- read the InstructGPT model card and evaluation methodology
- study labeling instructions used for human feedback data collection
- compare GPT-3 and InstructGPT outputs on NLP benchmarks
- learn how supervised fine-tuning plus reward modeling reduces toxicity
- research alignment techniques for large language models

## When to choose
- you want the authoritative artifacts and documentation behind the InstructGPT paper
- you are studying RLHF and model alignment from primary sources
- you need the official model card or labeling guidelines for citation or replication

## When to avoid
- you need runnable RLHF training code - this repo ships no implementation
- you want the actual InstructGPT weights or datasets - they are not included
- you need a maintained library - the repo is a static paper companion

## Facets
- artifact type: learning-resource
- maturity: maintenance
- function: machine-learning, llm-training, reinforcement-learning, nlp
- domain: large-language-models, machine-learning, artificial-intelligence, tutorials
- platform: python
- tags: rlhf, instructgpt, human-feedback, model-alignment, paper-companion, model-card, gpt-3, fine-tuning, natural-language-processing, gpu, linux

## Member repositories
- openai/following-instructions-human-feedback (main) score 10

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:09.716827+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T05:05:31.638080+00:00, confidence not recorded.
  - readme: https://github.com/openai/following-instructions-human-feedback (fetched 2026-08-28T04:04:09.716827+00:00, sha 8ba860f7ffca)
- Data as of 2026-08-30T08:39:29.467469+00:00.
