# verazuo/jailbreak_llms

[CCS'24] A dataset consists of 15,140 ChatGPT prompts from Reddit, Discord, websites, and open-source datasets (including 1,405 jailbreak prompts).

Repository: https://github.com/verazuo/jailbreak_llms
Canonical: https://ross.abutalabs.com/products/jailbreak_llms
Homepage: https://jailbreak-llms.xinyueshen.me/
Language: Jupyter Notebook
License: MIT
License Family: permissive
Topics: chatgpt, jailbreak, llm, prompt, large-language-model, llm-security, jailbreaking
Last push: 2024-12-24T08:30:58+00:00

## Health v2 (maintenance only)
Score: 28/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 80
- inputs: {"age_days": 1128, "days_push": 617, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3791, forks 329 (observed 2026-08-28T04:08:18.798204+00:00)

## What it is
A research dataset accompanying the ACM CCS 2024 paper 'Do Anything Now', containing 15,140 ChatGPT prompts collected from Reddit, Discord, prompt-aggregation websites, and open-source datasets between December 2022 and December 2023, of which 1,405 are identified jailbreak prompts. It supports the JailbreakHub measurement framework, which analyzes jailbreak communities and attack strategies such as prompt injection and privilege escalation, and evaluates safeguard failure rates across six popular LLMs.

## Use cases
- find a dataset of real jailbreak prompts targeting ChatGPT
- evaluate how well LLM safeguards resist adversarial prompts
- study how jailbreak prompts spread across Reddit and Discord
- train a classifier to detect jailbreak or harmful prompts
- red team an LLM using documented in-the-wild attack prompts
- benchmark multiple LLMs against prompt injection attacks
- research the accounts and communities that create jailbreak prompts

## When to choose
- You need a large, labeled corpus of in-the-wild jailbreak and normal prompts for LLM safety research
- You want peer-reviewed data (CCS 2024) with measured attack success rates to benchmark model defenses
- You need to study the evolution, strategies, and authorship of jailbreak prompts over a full year

## When to avoid
- You need prompts collected after December 2023, since the dataset's collection window has closed
- You want an attack tool or defense framework rather than a dataset for analysis
- Your project cannot handle harmful text, as the data contains offensive and dangerous content

## Facets
- artifact type: dataset
- maturity: stable
- function: security, machine-learning, nlp, benchmarking, testing
- domain: artificial-intelligence, large-language-models, security
- platform: python, cross-platform
- tags: llm-security, jailbreak-prompts, adversarial-prompts, ai-safety, red-teaming, prompt-injection, chatgpt, research-dataset, huggingface-dataset, llm-evaluation, ccs-2024, content-safety, natural-language-processing

## Member repositories
- verazuo/jailbreak_llms (main) score 28

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:18.798204+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:27:26.515081+00:00, confidence not recorded.
  - readme: https://github.com/verazuo/jailbreak_llms (fetched 2026-08-28T04:08:18.798204+00:00, sha 5fa46012b1d5)
  - homepage: https://jailbreak-llms.xinyueshen.me/ (fetched 2026-08-29T09:22:11.370921+00:00, sha da6bd581beb5)
- Data as of 2026-08-30T08:39:29.467469+00:00.
