Independent reference. No vendor affiliation. Reprinted scores dated to their source; elsewhere we link the official board. Editorial policy.
Abstract
What214 realistic web-browsing tasks, 5 to 90 minutes of human effort each.
WhoYoran, Amouyal, Malaviya, Bogin, Press, Berant (Tel Aviv University, EMNLP 2024).
BaselinesBrowser agent 25.2%, closed-book 11.1% in the launch paper; far from saturated.
Projectassistantbench.github.io
Section II.v Agent Benchmarks|Reviewed 2026

AssistantBench: 214 Realistic Web Tasks for Assistant Agents

The benchmark designed around what users actually ask assistants to do: multi-step web research with a verifiable answer.

I

Construction

The Tel Aviv University team built AssistantBench by sourcing tasks from three streams: (1) MTurk worker submissions of recent tasks they actually wanted help with, (2) university administrator workflows (course catalogue lookups, scholarship eligibility checks), and (3) author-curated tasks designed to stress specific failure modes. Each task ships with a gold answer extracted by humans and a verification protocol.

The benchmark deliberately excludes tasks that are answerable from training data alone. Closed-book GPT-4 hits only 11.1% on the test set, which the authors use to argue that browsing capability, not memory, is what is being measured.

II

Reading the scores

The launch paper's anchors are clear: the strongest browser agent reached 25.2%, a baseline web agent 14.4%, and browsing-disabled closed-book GPT-4 only 11.1%, which confirms the tasks require live web research rather than recall. We do not reprint a per-model table of current scores, because a responsibly-sourced one cannot be published here: there is no continuously-maintained board we re-verify, current numbers come from vendor reports under different browser harnesses, and the live-web design means a task's gold answer can change between runs, so 2024 and 2026 numbers are not strictly comparable.

For current results, read the official AssistantBench project page and each submission's harness together. To choose the right benchmark for your use case rather than chase a single figure, start from the homepage task picker.

III

Limitations

Live-web evaluation has a reproducibility cost. A 2024 task that asked for the rent of a specific listing has a different gold answer in 2026 because the listing changed. The AssistantBench team partly addresses this by snapshotting page contents at task creation time, but agents in the live browser see the current page, not the snapshot. This means scores from 2024 papers and 2026 papers are not strictly comparable for time-sensitive tasks.

GAIA: the other realistic assistant benchmarkMind2Web for action predictionBrowser-agent benchmarks compared
Reader Questions
Q.01What does AssistantBench test?+
AssistantBench is 214 web-browsing tasks designed to take a human between 5 and 90 minutes. Examples include 'find the average price of a 2-bedroom rental in three named neighbourhoods' or 'list the speakers at the most recent NeurIPS keynote and their affiliations'. Tasks require multi-page navigation, cross-source synthesis, and structured-output extraction.
Q.02What is the headline score?+
In the launch paper, the strongest browser agent scored 25.2% accuracy, a baseline web agent from the same group scored 14.4%, and closed-book browsing-disabled GPT-4 scored only 11.1%, which establishes that the questions are not solvable from memory alone. Those are the cleanly-sourced anchors. We do not reprint a per-model table of current scores: there is no continuously-maintained board we re-verify, vendor numbers use different browser harnesses, and live-web drift means a 2024 task and its 2026 re-run are not strictly comparable.
Q.03How does AssistantBench score answers?+
Most answers are short structured outputs (numbers, lists, JSON-like records). The scoring function is a programmatic match with normalisation for units, ordering, and minor formatting. Tasks where the answer is a free-text summary are excluded from the scored split because LLM-as-judge introduces variance the paper authors wanted to avoid.
Q.04Why is GPT-4 stuck below 30%?+
The failure analysis identified three dominant causes: (1) the agent stops too early and submits a partial answer; (2) the agent fails to verify an answer against a second source and accepts the first plausible value; (3) the agent's JSON extraction breaks on complex web layouts. None of these are model capability ceilings, they are harness limitations, which is why scores rise quickly when better harnesses ship.
Q.05How is this different from WebArena and Mind2Web?+
WebArena uses controlled simulated environments (a self-hosted Reddit, a self-hosted Gitea). Mind2Web uses static page snapshots with offline action prediction. AssistantBench uses the live public web. The cost is that the web changes under the agent and reproducibility weakens over time; the benefit is that AssistantBench measures what users actually want.

Sources

  1. [1] Yoran et al. (2024): arxiv.org/abs/2407.15711
  2. [2] AssistantBench project: assistantbench.github.io
From the editor

Benchmarking Agents Review is published by Digital Signet, an independent firm that builds and ships AI agents in production for mid-market companies. If you are evaluating, designing, or productionising LLM agents and want a working second opinion, get in touch.

Book a 30-min scoping callDigital Signet →

30 minutes, free, independent.·1-page action plan within 48h.·Honest if not the right fit.