PKU-Alignment/safe-rlhf
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback observed · 2026-08-28
Health v2 · maintenance only
53/100
- Activity 53
- Release rhythm 35
- Longevity 86
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1206
- days_rel: n/a
- days_push: 282
- n_releases_24m: 0
Adoption not part of the score
1611 stars · 133 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Beaver is a modular open-source framework from Peking University for training large language models with SFT, RLHF, and Safe RLHF (constrained reward maximization). It includes a large human-labeled preference dataset (up to 1M pairs) and pre-trained reward/cost model checkpoints.
Use cases
- train an LLM with RLHF
- align a language model with safety constraints
- run supervised fine-tuning on LLaMA or OPT
- train a reward model from human preferences
- train a cost model for harmlessness
- research safe reinforcement learning from human feedback
- evaluate LLM safety with BIG-bench or GPT-4 evaluation
When to choose
- you need a reproducible RLHF or Safe RLHF training pipeline
- you want human-labeled helpfulness and harmlessness preference data
- you are doing alignment research with reward and cost models
- you want pre-trained aligned model checkpoints like Beaver-7B
When to avoid
- you only need to run inference on an existing LLM
- you lack multi-GPU resources for large model training
- you need a general-purpose RL library unrelated to LLM alignment
Facets
framework · maturity active
llm-training machine-learning deep-learning rag large-language-models machine-learning artificial-intelligence deep-learning python rlhf safe-rlhf alignment ai-safety reward-model cost-model sft deepspeed preference-learning llm-alignment gpu linux
2 sources
- readme: https://github.com/PKU-Alignment/safe-rlhf · fetched 2026-08-28 · 0f1ee9303a3c
- homepage: https://pku-beaver.github.io · fetched 2026-08-29 · df9592ac827c
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| PKU-Alignment/safe-rlhf | main | 53 |
For agents
markdown · JSON · MCP: product_card(name="PKU-Alignment/safe-rlhf")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem