# tau-bench

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Repository: https://github.com/sierra-research/tau2-bench
Canonical: https://ross.abutalabs.com/products/tau-bench
Homepage: https://www.taubench.com
Language: Python
License: MIT
License Family: permissive
Topics: benchmark, llm, ai, language-model-agent, conversational-agents
Last push: 2026-08-18T17:07:54+00:00
Link (homepage): https://www.taubench.com

## Health v2 (maintenance only)
Score: 79/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 98, release rhythm 82, longevity 32
- inputs: {"age_days": 450, "days_push": 15, "days_rel": 42, "gap_med": 83.0, "n_releases_24m": 5}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1882, forks 472 (observed 2026-08-28T04:05:48.658890+00:00)

## What it is
τ-Bench (tau2-bench) is a Python benchmark for evaluating LLM agents on tool-agent-user interaction in real-world domains like retail, airline, telecom, and banking. It scores agents on conversations with simulated users, tool calls, knowledge retrieval, and policy adherence, including voice full-duplex evaluation, with a live public leaderboard.

## Use cases
- benchmark my llm agent on tool calling and customer service tasks
- evaluate how reliably an agent follows domain policies
- compare models on multi-turn conversational agent tasks
- test voice agents on real-time full-duplex conversations
- evaluate rag pipelines on knowledge-intensive banking tasks
- measure pass^k reliability of agent trajectories
- reproduce published agent benchmark scores

## When to choose
- you need standardized, verifiable evaluation of tool-using conversational agents
- you want to compare your model against a public leaderboard
- you need both text and voice agent evaluation with realistic simulated users

## When to avoid
- you need a general-purpose agent framework for building production agents rather than evaluating them
- your domain is unrelated to the provided retail/airline/telecom/banking scenarios and you cannot author custom tasks
- you need cheap, fast evals - simulated user LLM calls add cost and latency

## Facets
- artifact type: dataset
- maturity: active
- function: benchmarking, agent-framework, llm-inference, rag, speech-recognition, testing
- domain: artificial-intelligence, large-language-models, chatbots, developer-tools, tutorials
- platform: python, cli, cross-platform
- tags: llm-benchmark, tool-use, agent-evaluation, conversational-agents, voice-agents, leaderboard, pass-k-metric, customer-service-domains, ai-agents, natural-language-processing

## Member repositories
- sierra-research/tau2-bench (main) score 79
- sierra-research/tau-bench (mirror) score 56

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:48.658890+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:13:44.976342+00:00, confidence not recorded.
  - readme: https://github.com/sierra-research/tau2-bench (fetched 2026-08-28T04:05:48.658890+00:00, sha d185613c5b82)
  - homepage: https://www.taubench.com (fetched 2026-08-29T10:53:06.274282+00:00, sha d9d1f969d5d7)
- Data as of 2026-08-30T08:39:29.467469+00:00.
