tau-bench resource
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains observed · 2026-08-28
Health v2 · maintenance only
79/100
- Activity 98
- Release rhythm 82
- Longevity 32
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 83.0
- age_days: 450
- days_rel: 42
- days_push: 15
- n_releases_24m: 5
Adoption not part of the score
1882 stars · 472 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
τ-Bench (tau2-bench) is a Python benchmark for evaluating LLM agents on tool-agent-user interaction in real-world domains like retail, airline, telecom, and banking. It scores agents on conversations with simulated users, tool calls, knowledge retrieval, and policy adherence, including voice full-duplex evaluation, with a live public leaderboard.
Use cases
- benchmark my llm agent on tool calling and customer service tasks
- evaluate how reliably an agent follows domain policies
- compare models on multi-turn conversational agent tasks
- test voice agents on real-time full-duplex conversations
- evaluate rag pipelines on knowledge-intensive banking tasks
- measure pass^k reliability of agent trajectories
- reproduce published agent benchmark scores
When to choose
- you need standardized, verifiable evaluation of tool-using conversational agents
- you want to compare your model against a public leaderboard
- you need both text and voice agent evaluation with realistic simulated users
When to avoid
- you need a general-purpose agent framework for building production agents rather than evaluating them
- your domain is unrelated to the provided retail/airline/telecom/banking scenarios and you cannot author custom tasks
- you need cheap, fast evals - simulated user LLM calls add cost and latency
Facets
dataset · maturity active
benchmarking agent-framework llm-inference rag speech-recognition testing artificial-intelligence large-language-models chatbots developer-tools tutorials python cli cross-platform llm-benchmark tool-use agent-evaluation conversational-agents voice-agents leaderboard pass-k-metric customer-service-domains ai-agents natural-language-processing
2 sources
- readme: https://github.com/sierra-research/tau2-bench · fetched 2026-08-28 · d185613c5b82
- homepage: https://www.taubench.com · fetched 2026-08-29 · d9d1f969d5d7
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| sierra-research/tau2-bench | main | 79 |
| sierra-research/tau-bench | mirror | 56 |
For agents
markdown · JSON · MCP: product_card(name="sierra-research/tau2-bench")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem