Ross ROSS = Recommend OSS · open-source software intelligence for agents

huggingface/alignment-handbook

Robust recipes to align language models with human and AI preferences observed · 2026-08-28

github.com/huggingface/alignment-handbook · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

66/100

  • Activity 84
  • Release rhythm 35
  • Longevity 78

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1104
  • days_rel: n/a
  • days_push: 99
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

5671 stars · 490 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A collection of robust training recipes and scripts from Hugging Face for aligning large language models with human and AI preferences, covering SFT, DPO, ORPO, and RLHF pipelines. It includes reproducible recipes behind models like Zephyr and SmolLM, along with associated datasets and evaluation guidance.

Use cases

  • fine-tune an LLM with DPO
  • run supervised fine-tuning on a chat dataset
  • align a model with human preferences using RLHF
  • reproduce the Zephyr training recipe
  • train a small instruct model like SmolLM
  • compare preference alignment methods like DPO vs KTO vs IPO

When to choose

  • you want battle-tested, reproducible recipes for LLM post-training
  • you're using the Hugging Face Transformers/TRL ecosystem
  • you need proven recipes for SFT plus preference alignment (DPO, ORPO, RLHF)

When to avoid

  • you only need inference or serving of LLMs
  • you want a general-purpose training framework rather than opinionated recipes
  • you're training non-language models

Facets

library · maturity active

llm-training machine-learning deep-learning large-language-models machine-learning deep-learning artificial-intelligence python rlhf dpo sft orpo fine-tuning transformers recipes zephyr gpu linux

7 sources

Member repositories

RepositoryRoleHealth v2
huggingface/alignment-handbookmain66

For agents

markdown · JSON · MCP: product_card(name="huggingface/alignment-handbook")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem