Accio-org/CommerceAgentBench resource
CommerceAgentBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services observed · 2026-08-28
Health v2 · maintenance only
57/100
- Activity 99
- Release rhythm 35
- Longevity 2
Flags: no_releases young
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 31
- days_rel: n/a
- days_push: 9
- n_releases_24m: 0
Adoption not part of the score
1192 stars · 85 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
CommerceAgentBench is a benchmark of 107 long-horizon agent tasks evaluated in high-fidelity, stateful, reproducible replicas of real online services (Shopify, Gmail, Stripe, Jira, etc.). It ships containerized task environments with deterministic or LLM-assisted verifiers and a public leaderboard for comparing model-harness pairs.
Use cases
- benchmark llm agents on real-world commerce workflows
- evaluate long-horizon agent task completion
- compare models on browser, CLI, and API/MCP tasks
- test agents against stateful replicas of real web services
- measure agent reliability with reproducible verifiers
- run cross-harness agent evaluations
When to choose
- you need reproducible, stateful evaluation of agents on multi-step business workflows
- you want to compare LLMs on browser, CLI, and API/MCP task execution
- you need deterministic or LLM-assisted grading of agent outcomes
When to avoid
- you need a lightweight single-turn QA benchmark
- your domain is unrelated to commerce or business workflows
- you want a benchmark that runs without container infrastructure
Facets
dataset · maturity active
benchmarking agent-framework testing mcp artificial-intelligence e-commerce large-language-models developer-tools python cli agent-evaluation llm-benchmark stateful-mocks long-horizon-tasks reproducibility leaderboard commerce-workflows containerized-eval ai-agents docker web-server
2 sources
- readme: https://github.com/Accio-org/CommerceAgentBench · fetched 2026-08-28 · 65769d002167
- homepage: https://commerce-agent-bench.site.accio.ai/ · fetched 2026-08-29 · fd4e8d205cec
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| Accio-org/CommerceAgentBench | main | 57 |
For agents
markdown · JSON · MCP: product_card(name="Accio-org/CommerceAgentBench")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem