Ross ROSS = Recommend OSS · open-source software intelligence for agents

llm-attacks/llm-attacks

Universal and Transferable Attacks on Aligned Language Models observed · 2026-08-28

github.com/llm-attacks/llm-attacks · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

28/100

  • Activity 0
  • Release rhythm 35
  • Longevity 81

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1134
  • days_rel: n/a
  • days_push: 761
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

4769 stars · 635 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Official research code for 'Universal and Transferable Adversarial Attacks on Aligned Language Models', implementing the GCG algorithm for automatically generating adversarial suffixes that jailbreak aligned LLMs. It targets open-source models like Vicuna and LLaMA-2 and demonstrates transfer of attacks to closed-source chatbots.

Use cases

  • generate adversarial suffixes that bypass LLM safety alignment
  • reproduce GCG jailbreak experiments on LLaMA-2 and Vicuna
  • study transferability of adversarial attacks to closed-source chatbots
  • evaluate robustness of aligned language models
  • run automated red-teaming experiments on LLMs

When to choose

  • you are an AI safety researcher studying LLM vulnerabilities
  • you need the reference implementation of the GCG attack algorithm
  • you want to reproduce the paper's experiments on open-weight models

When to avoid

  • you need a production-ready LLM security tool
  • you want a maintained pip-installable library (use nanoGCG instead)
  • you lack GPU resources to run 7B-parameter models

Facets

library · maturity maintenance

machine-learning llm-training security penetration-testing large-language-models security artificial-intelligence machine-learning python adversarial-attacks jailbreak llm-safety gcg alignment research-code linux gpu

2 sources

Member repositories

RepositoryRoleHealth v2
llm-attacks/llm-attacksmain28

For agents

markdown · JSON · MCP: product_card(name="llm-attacks/llm-attacks")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem