OpenGenerativeAI/llm-colosseum
Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM observed · 2026-08-28
Health v2 · maintenance only
28/100
- Activity 12
- Release rhythm 28
- Longevity 63
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 41
- age_days: 893
- days_rel: 651
- days_push: 530
- n_releases_24m: 2
Adoption not part of the score
1484 stars · 182 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
LLM Colosseum is a benchmarking tool that pits large language models against each other in real-time Street Fighter III matches to evaluate their decision-making quality. It produces an ELO-based leaderboard from hundreds of automated fights between models like GPT-4o, Pixtral, Claude, and Llama.
Use cases
- benchmark llms by playing street fighter
- compare llm decision making in real time
- elo ranking of large language models
- evaluate llm agents in a game environment
- test which ai model is the best fighter
- alternative llm benchmark beyond static tests
When to choose
- you want a novel, real-time evaluation of LLM reasoning and adaptability
- you want to compare vision and text modes of frontier models in an interactive setting
- you want an entertaining demo of LLM agent capabilities
When to avoid
- you need rigorous, reproducible academic benchmarks
- you need standard NLP task evaluation like MMLU or code generation
- you need a production LLM evaluation pipeline
Facets
application · maturity active
benchmarking llm-inference agent-framework machine-learning large-language-models artificial-intelligence gaming-tools analytics python browser llm-evaluation elo-ranking street-fighter game-benchmark reinforcement-learning-comparison jupyter-notebook web desktop
2 sources
- readme: https://github.com/OpenGenerativeAI/llm-colosseum · fetched 2026-08-28 · f830e0bc546d
- homepage: https://llm-colosseum.phospho.ai · fetched 2026-08-29 · a6873599d06f
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| OpenGenerativeAI/llm-colosseum | main | 28 |
For agents
markdown · JSON · MCP: product_card(name="OpenGenerativeAI/llm-colosseum")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem