Independent reference. No vendor affiliation. Reprinted scores dated to their source; elsewhere we link the official board. Editorial policy.
Abstract
What165 tasks (115 retail, 50 airline) across two simulated customer-service domains; policy-grounded dialogue with tool calls.
WhoSierra Research: Yao, Shinn, Razavi, Narasimhan.
Board topClaude 3.5 Sonnet (Oct 2024) pass^1 retail 69.2% / airline 46.0%; GPT-4o 60.4% / 42.0%. Original board frozen at late-2024 models.
Paperarxiv.org/abs/2406.12045
Section II.ix Agent Benchmarks|Reviewed 2026|Task counts (retail 115, airline 50 = 165) corrected and board figures re-verified 10 Jul 2026; Gemini 2.5 Pro HAL airline 'clean' 71.8% re-verified 10 Jul 2026

Tau-Bench Retail and Airline: 165 Tasks, Frontier Pass^1 Below 70%

Two customer-service simulations that surface the consistency tax most benchmarks hide.

I

The Domains

Retail covers a typical consumer e-commerce help desk: order modification, returns, refunds, exchanges, shipping changes. The agent reads a written policy (retail_policy.md is part of the benchmark and is shown to the agent in-context) and must follow it. The simulated user is itself an LLM playing a customer with a goal.

Airline is a tougher version of the same shape. The policies are stricter: refund eligibility depends on fare class, change fees apply, some changes are forbidden, and the agent must look up loyalty-program rules. The action space is larger and the cost of a wrong tool call is higher.

II

SOTA Progression

Date
Tier / Score
Note
Jun 2024
GPT-4o pass^1 Retail 60.4%, Airline 42.0%
Official board figures; the launch paper (arXiv:2406.12045) originally reported airline at 35.2%.
Oct 2024
Claude 3.5 Sonnet (20241022) pass^1 Retail 69.2%, Airline 46.0%
Top of the official board; the highest pass^1 it records.
2025-26
Original board frozen at the late-2024 model set
Repo states its tasks are not updated; newer models go to the successor tau2-bench / tau3-bench, not this board.
III

The pass^k Insight

Most published agent numbers are best-of-k (try k times, take the best). Tau-Bench refuses that pattern. pass^k requires all k runs to succeed, which surfaces consistency failures that best-of-k hides. In the original paper GPT-4o's retail score falls from about 60% at pass^1 toward roughly 25% by pass^8, a consistency collapse best-of-k scoring would hide. That gap is the real metric for "would I ship this in production?" because production users do not retry.

Tau-Bench overviewBFCL for atomic function-callingTool-use benchmarks compared
Reader Questions
Q.01What does the Tau-Bench Retail domain test?+
Tau-Bench Retail (Sierra Research, 2024) is a 115-task simulated retail customer-service environment. The agent talks to a simulated user, must follow company policies on returns, exchanges, refunds, and order modifications, and must call tools (cancel order, issue refund, modify shipping). A task is scored 1 if the final database state matches what policy required, 0 otherwise.
Q.02What does the Airline domain test?+
Tau-Bench Airline is a 50-task simulated airline customer-service environment. Tasks include booking, change-of-itinerary, cancellation, special-meal requests, and frequent-flyer-program lookups. The policies are stricter than Retail (refunds depend on fare class, change fees apply, some changes are forbidden), which makes Airline harder than Retail for most models.
Q.03What is pass^k and why is it different from pass@k?+
pass^k measures consistency: the same task is run k times, and pass^k is the fraction of tasks where the agent succeeds on all k runs. pass@k (HumanEval style) is the fraction of tasks where the agent succeeds at least once across k runs. pass^k is harsher because it requires reproducible success. For agentic systems, pass^k is the more honest signal because users do not retry until something works.
Q.04What are the headline numbers on the official board?+
The official sierra-research/tau-bench board tops out at Claude 3.5 Sonnet (20241022): pass^1 retail 69.2%, airline 46.0%. GPT-4o sits at retail 60.4%, airline 42.0%. That board is frozen at the late-2024 model set, so newer models (Gemini 2.5 Pro, Claude 4.x, GPT-5) are not on it: their Tau-Bench numbers come from vendor self-reports or the successor tau2-bench / tau3-bench and are not directly comparable. The retail-airline gap is the consistency story: airline's stricter policy surface punishes any tool-call mistake.
Q.05What does Gemini 2.5 Pro score on Tau-Bench retail?+
There is no Gemini 2.5 Pro figure on the original sierra-research/tau-bench board, which is frozen at the late-2024 model set (its top entry is Claude 3.5 Sonnet at retail 69.2%, airline 46.0%). Numbers for newer models come from independent reproductions or the successor tau2-bench, and are harness-dependent rather than directly comparable to that frozen board. Princeton's Holistic Agent Leaderboard (HAL) reports Gemini 2.5 Pro at 71.8% on its 'clean' variant of tau-bench airline but does not currently publish a retail figure for it. So treat any single quoted retail percentage as specific to its tau-bench version (original v1, a 'clean' re-run, or tau2-bench) and confirm which one the source used before comparing models.
Q.06How is Tau-Bench different from BFCL?+
BFCL evaluates atomic function-calling: pick the right tool, fill the right arguments, in isolation. Tau-Bench evaluates policy-grounded multi-turn dialogue where the agent must combine tool calls with dialogue management, policy adherence, and state tracking. Tau-Bench is a strict superset task. Strong BFCL scores are necessary but not sufficient for strong Tau-Bench scores.

Sources

  1. [1] Yao et al. (2024): arxiv.org/abs/2406.12045
  2. [2] Tau-Bench repository (pass^1 board figures; README notes the original tasks are frozen): github.com/sierra-research/tau-bench. Accessed 17 Jun 2026.
  3. [3] Successor benchmark for current models: github.com/sierra-research/tau2-bench (tau2-bench / tau3-bench).
  4. [4] Independent reproductions for current models (Gemini 2.5 Pro tau-bench airline 'clean' 71.8%): Princeton Holistic Agent Leaderboard (HAL) reliability dashboard, hal.cs.princeton.edu/reliability. Accessed 25 Jun 2026.
  5. [5] Sierra announcement: sierra.ai/blog/benchmarking-ai-agents
From the editor

Benchmarking Agents Review is published by Digital Signet, an independent firm that builds and ships AI agents in production for mid-market companies. If you are evaluating, designing, or productionising LLM agents and want a working second opinion, get in touch.

Book a 30-min scoping callDigital Signet →

30 minutes, free, independent.·1-page action plan within 48h.·Honest if not the right fit.