Independent reference. No vendor affiliation. Reprinted scores dated to their source; elsewhere we link the official board. Editorial policy.
Colophon|Last verified June 2026

About benchmarkingagents.com

An independent reference for AI agent and LLM benchmarks. Built by Digital Signet. No vendor affiliation, no paid placements, no newsletter capture.

What this site is

benchmarkingagents.com is an independent reference for the public benchmarks used to evaluate large language models and AI agents in 2026. Coverage spans the 17 most-cited benchmarks across knowledge (MMLU, MMLU-Pro, MMMU, Humanity's Last Exam), coding (HumanEval, MBPP, LiveCodeBench, SWE-bench Verified), reasoning (GPQA, GPQA-Diamond, ARC-AGI, ARC-AGI-2, BIG-Bench Hard), agentic capability (SWE-bench Verified, WebArena, AgentBench, OSWorld, Terminal-Bench, Tau-Bench), multimodal (MMMU, MathVista, ChartQA), and human preference (LMSYS Chatbot Arena, MT-Bench).

It also covers the practitioner stack that production teams actually use to evaluate their own systems: custom golden datasets, LLM-as-judge methodology, RAG evaluation frameworks like Ragas, online production monitoring, and the eight evaluation platforms in widest use (Braintrust, Langfuse, LangSmith, Arize Phoenix, DeepEval, Patronus, Helicone, PromptLayer).

It is not a leaderboard that scrapes scores without context. Where a benchmark has a current, checkable primary board, the homepage picker reprints its top rows with the date verified, the source, and the caveats (scaffold, evaluation mode, saturation) that any quoted score depends on; where no trustworthy board exists, it says so rather than inventing a table, and the benchmark page links the official leaderboard for the live number. Where a benchmark has saturated or has documented contamination risk, the page says so explicitly.

Why this site exists

The benchmark landscape has three structural problems that most AI coverage ignores or treats as footnotes. First, contamination. MMLU test questions appear verbatim in Common Crawl, the primary web-crawl dataset behind most pre-training corpora. HumanEval problems are near-duplicates of LeetCode solutions that pervade GitHub and Stack Overflow crawls. SWE-bench issues have solutions in the same repositories' public git history. A high score on a contaminated benchmark may reflect memorisation as much as reasoning, and the leaderboard cannot tell you which.

Second, saturation. MMLU saturated in 2024. HumanEval, HellaSwag and WinoGrande saturated in 2023 or 2024. All frontier models now score within statistical noise of each other on these benchmarks. The field moved to MMLU-Pro (10-choice, CoT-required), GPQA-Diamond, ARC-AGI-2 and LiveCodeBench specifically because the previous generation no longer discriminates. Most comparison sites still quote the saturated versions, because the numbers are larger and more recognisable.

Third, self-reporting. Model cards are published by the same company that built the model. Common issues: publishing only benchmarks where the model performs well, omitting evaluation setup (N-shot, CoT, harness, test set version), and reporting scores under maximally favourable conditions that independent replications struggle to reproduce. The benchmark-literate reader treats vendor-published scores as a starting point, not a final answer.

This site exists to apply a consistent six-question rubric to every benchmark score it cites, document the methodology behind every number, and refuse to publish vendor-favourable shortcuts when the underlying signal is noisy.

Who builds this

Oliver Wakefield-Smith
Editor

Oliver Wakefield-Smith

Founder, Digital Signet

benchmarkingagents.com is one of a small cluster of Digital Signet reference sites covering the cost, capability and evaluation surface of frontier AI. Sister sites focus on per-token model pricing; this one covers what the resulting models can and cannot do, and how to measure it without being captured by vendor framing.

Sister sites in the Digital Signet AI-pricing cluster
claudeapipricing.com

Independent reference for Claude API token pricing across models, batch tier, prompt caching.

embeddingcost.com

Multi-provider AI embedding pricing, vector DB storage cost, RAG scenarios.

geminipricing.com

Google Gemini API pricing reference; cross-checks Vertex AI surface.

contextcost.com

Per-million-token cost calculator across model providers; latency and cost trade-offs.

Editorial position

Independent reference. No vendor affiliation. No paid placements on the tool-comparison pages. No sponsored entries in any benchmark or evaluation framework cited. Provider order in tables is determined alphabetically or by category, not by any commercial relationship.

Outbound links to evaluation tool vendors on the /tools-compared page go directly to each vendor and are not monetised today. If an affiliate link is ever added it will be labelled, and it would not influence the order tools appear, the rubric used to compare them, or which tools are recommended for which use case. The neutral editorial framing holds regardless.

No display advertising. No newsletter capture. No content sponsorships. No vendor lead-routing. The site is a reference, not a lead-generation funnel.

What this site covers

Editorial principles

Principle

Source pattern

Where a benchmark has a current, checkable primary leaderboard, the homepage picker reprints its top rows fetched from that source on a named verification date, with a caveat. Where the source publishes no per-model board, or its board has not kept up with the current frontier, the site says so rather than invent rows, and individual benchmark pages link the official board (swebench.com, arcprize.org, lmsys.org, Papers With Code) rather than freeze a number. Where vendor cards and independent leaderboards disagree, the more independent source is preferred.

Principle

Dating and saturation

Frontier benchmarks move fast: SOTA in November is not SOTA in May. Every reprinted row carries the date it was verified against the primary source; a benchmark whose official board has not kept up with the current frontier is shown with that caveat, and once a benchmark crosses the 90 percent saturation threshold the page flags it explicitly rather than continuing to report the highest reported number.

Principle

N-shot, CoT and harness disclosure

A 5-shot CoT score and a 0-shot greedy score are not the same number, so a bare percentage is close to meaningless. The benchmark pages explain which evaluation settings matter for each benchmark and why a reprinted frontier score without its harness is not comparable; for the live number under a disclosed setup the site links the official leaderboard.

Principle

Saturation flagged

MMLU, HumanEval, MBPP, HellaSwag and WinoGrande are saturated across frontier models in 2026. The site states this on the benchmark page and points readers to current benchmarks (MMLU-Pro, GPQA-Diamond, ARC-AGI-2, LiveCodeBench, SWE-bench Verified) for live frontier comparison.

Principle

Contamination flagged

Public benchmarks with high contamination risk (MMLU, HumanEval, HellaSwag) are flagged on the relevant pages. Independent verification of vendor-claimed scores on contaminated benchmarks is impossible because the contamination is structural, not a matter of bad actors.

Principle

No vendor ranking shortcuts

The site does not produce a single composite leaderboard that purports to rank all frontier models against each other. Different benchmarks measure different things. A composite ranking would hide the trade-offs that matter for production model selection.

Refresh cadence

SOTA benchmark scores are re-verified against the official leaderboard for each benchmark, prioritised by search position: the homepage task-to-benchmark picker and the agent-benchmark cluster heads (SWE-bench Verified, WebArena, AgentBench) first, the long tail after. The most recent sweep, covering the picker's six cited sources, closed in June 2026. Where a cited leaderboard turns out not to publish per-model rows, the page says so rather than reprinting numbers we cannot trace.

The verification date is held in a single constant (LAST_VERIFIED_DATE) in src/lib/schema.ts. Footer text, schema dateModified, and visible headings all read from that single source. This is a deliberate design choice so cosmetic-refresh leaks (rolling a date forward without doing the underlying verification work) are structurally prevented.

Disclosures

  • 01.No affiliate parameters on benchmark or vendor URLs cited in editorial content.
  • 02.Outbound links to evaluation-tool vendors on the /tools-compared page go directly to each vendor and are not monetised today. Any affiliate link added later will be labelled; tool order is alphabetical by category and would not be influenced by it.
  • 03.Not affiliated with Anthropic, OpenAI, Google, Meta, Mistral, Cohere, AWS Bedrock, Azure OpenAI, or any other listed model vendor.
  • 04.Not affiliated with Braintrust, Langfuse, LangSmith, Arize, DeepEval, Patronus, Helicone, PromptLayer, or any other listed eval-tool vendor.
  • 05.Not affiliated with Princeton, Stanford, OpenAI, Allen AI, EleutherAI, LMSYS, HuggingFace, or any benchmark-producing organisation.

Contact and corrections

Corrections welcome. If you find a misquoted figure, a stale date, a methodology error, or a citation that does not resolve to the claimed source, email editor@benchmarkingagents.com and the correction will land in the next refresh pass.

Five business days is the target response window. For substantive corrections (not just a typo) the corrected page also adds a short corrigendum note at the bottom and rolls the dateModified forward in the schema. Cosmetic typos roll silently.

For commercial enquiries (sponsored content, paid placements, lead-routing arrangements) the answer is no. The site is a reference, and accepting any of those would compromise the editorial position above.

From the editor

Benchmarking Agents Review is published by Digital Signet, an independent firm that builds and ships AI agents in production for mid-market companies. If you are evaluating, designing, or productionising LLM agents and want a working second opinion, get in touch.

Book a 30-min scoping callDigital Signet →

30 minutes, free, independent.·1-page action plan within 48h.·Honest if not the right fit.