Ross ROSS = Recommend OSS · open-source software intelligence for agents

OpenGenerativeAI/llm-colosseum

Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM observed · 2026-08-28

github.com/OpenGenerativeAI/llm-colosseum · homepage · Jupyter Notebook · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

28/100

  • Activity 12
  • Release rhythm 28
  • Longevity 63
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 41
  • age_days: 893
  • days_rel: 651
  • days_push: 530
  • n_releases_24m: 2

Full methodology

Adoption not part of the score

1484 stars · 182 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

LLM Colosseum is a benchmarking tool that pits large language models against each other in real-time Street Fighter III matches to evaluate their decision-making quality. It produces an ELO-based leaderboard from hundreds of automated fights between models like GPT-4o, Pixtral, Claude, and Llama.

Use cases

  • benchmark llms by playing street fighter
  • compare llm decision making in real time
  • elo ranking of large language models
  • evaluate llm agents in a game environment
  • test which ai model is the best fighter
  • alternative llm benchmark beyond static tests

When to choose

  • you want a novel, real-time evaluation of LLM reasoning and adaptability
  • you want to compare vision and text modes of frontier models in an interactive setting
  • you want an entertaining demo of LLM agent capabilities

When to avoid

  • you need rigorous, reproducible academic benchmarks
  • you need standard NLP task evaluation like MMLU or code generation
  • you need a production LLM evaluation pipeline

Facets

application · maturity active

benchmarking llm-inference agent-framework machine-learning large-language-models artificial-intelligence gaming-tools analytics python browser llm-evaluation elo-ranking street-fighter game-benchmark reinforcement-learning-comparison jupyter-notebook web desktop

2 sources

Member repositories

RepositoryRoleHealth v2
OpenGenerativeAI/llm-colosseummain28

For agents

markdown · JSON · MCP: product_card(name="OpenGenerativeAI/llm-colosseum")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem