Ross ROSS = Recommend OSS · open-source software intelligence for agents

microsoftarchive/promptbench

A unified evaluation framework for large language models observed · 2026-08-28

github.com/microsoftarchive/promptbench · homepage · Python · MIT (permissive) · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 68
  • Release rhythm 35
  • Longevity 84

Flags: no_releases archived

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1177
  • days_rel: n/a
  • days_push: 194
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2819 stars · 221 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

PromptBench is a unified Python library from Microsoft for evaluating and understanding large language models across many datasets, models, and prompt techniques. It includes adversarial prompt attack modules and robustness analysis to measure how LLM performance degrades under perturbed prompts.

Use cases

  • benchmark llm performance across multiple datasets
  • evaluate prompt robustness against adversarial attacks
  • compare gpt-4o, gemini, mistral and open-source models
  • measure how prompt wording affects model accuracy
  • run multi-prompt evaluation efficiently
  • evaluate multimodal llm capabilities

When to choose

  • you need a standardized harness to compare many LLMs on academic benchmarks
  • you want to study prompt sensitivity and adversarial robustness of LLMs
  • you need reproducible evaluation with support for major commercial and open models

When to avoid

  • you need a production serving or inference framework rather than evaluation
  • you want lightweight ad-hoc testing of a single model without benchmark overhead
  • you need a hosted leaderboard service rather than a local library

Facets

framework · maturity maintenance

benchmarking llm-inference prompt-engineering testing large-language-models machine-learning artificial-intelligence python cross-platform llm-evaluation adversarial-robustness prompt-robustness benchmarking-suite microsoft

1 source

Member repositories

RepositoryRoleHealth v2
microsoftarchive/promptbenchmain10

For agents

markdown · JSON · MCP: product_card(name="microsoftarchive/promptbench")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem