# centerforaisafety/HarmBench

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Repository: https://github.com/centerforaisafety/HarmBench
Canonical: https://ross.abutalabs.com/products/harmbench
Homepage: https://harmbench.org
Language: Jupyter Notebook
License: MIT
License Family: permissive
Last push: 2024-08-16T04:37:37+00:00

## Health v2 (maintenance only)
Score: 26/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 67
- inputs: {"age_days": 943, "days_push": 747, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1033, forks 157 (observed 2026-08-28T04:03:18.437814+00:00)

## What it is
HarmBench is a standardized, open-source evaluation framework for automated red teaming of large language models, comparing attack methods and defenses across many LLMs. It provides an evaluation pipeline, classifiers, precomputed test cases, and adversarial training code for improving LLM robustness.

## Use cases
- benchmark LLM jailbreak attacks against multiple models
- evaluate how robust an LLM is to red teaming methods
- compare automated red teaming techniques on a standard test suite
- train LLMs to refuse harmful requests via adversarial training
- run safety evaluations before deploying an LLM

## When to choose
- you need a standardized benchmark for LLM attack/defense evaluation
- you want to compare many red teaming methods or target LLMs reproducibly
- you are researching LLM safety and adversarial robustness

## When to avoid
- you need general-purpose ML benchmarking unrelated to LLM safety
- you want a lightweight safety filter for production rather than an evaluation harness
- your models are not LLMs

## Facets
- artifact type: framework
- maturity: active
- function: benchmarking, testing, llm-inference, machine-learning
- domain: artificial-intelligence, large-language-models, security, machine-learning
- platform: python, cli
- tags: llm-safety, red-teaming, adversarial-attacks, llm-evaluation, jailbreaks, ai-safety, gpu

## Member repositories
- centerforaisafety/HarmBench (main) score 26

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:18.437814+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:06:39.798996+00:00, confidence not recorded.
  - readme: https://github.com/centerforaisafety/HarmBench (fetched 2026-08-28T04:03:18.437814+00:00, sha 0995eb79bc8c)
  - homepage: https://harmbench.org (fetched 2026-08-29T13:06:37.835365+00:00, sha a27fcb336529)
- Data as of 2026-08-30T08:39:29.467469+00:00.
