Ross ROSS = Recommend OSS · open-source software intelligence for agents

centerforaisafety/HarmBench

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal observed · 2026-08-28

github.com/centerforaisafety/HarmBench · homepage · Jupyter Notebook · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

26/100

  • Activity 0
  • Release rhythm 35
  • Longevity 67

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 943
  • days_rel: n/a
  • days_push: 747
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1033 stars · 157 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

HarmBench is a standardized, open-source evaluation framework for automated red teaming of large language models, comparing attack methods and defenses across many LLMs. It provides an evaluation pipeline, classifiers, precomputed test cases, and adversarial training code for improving LLM robustness.

Use cases

  • benchmark LLM jailbreak attacks against multiple models
  • evaluate how robust an LLM is to red teaming methods
  • compare automated red teaming techniques on a standard test suite
  • train LLMs to refuse harmful requests via adversarial training
  • run safety evaluations before deploying an LLM

When to choose

  • you need a standardized benchmark for LLM attack/defense evaluation
  • you want to compare many red teaming methods or target LLMs reproducibly
  • you are researching LLM safety and adversarial robustness

When to avoid

  • you need general-purpose ML benchmarking unrelated to LLM safety
  • you want a lightweight safety filter for production rather than an evaluation harness
  • your models are not LLMs

Facets

framework · maturity active

benchmarking testing llm-inference machine-learning artificial-intelligence large-language-models security machine-learning python cli llm-safety red-teaming adversarial-attacks llm-evaluation jailbreaks ai-safety gpu

2 sources

Member repositories

RepositoryRoleHealth v2
centerforaisafety/HarmBenchmain26

For agents

markdown · JSON · MCP: product_card(name="centerforaisafety/HarmBench")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem