Toolathlon-Verified

Realistic multi-step chores that require an agent to juggle many real software tools over a long session, not just answer questions.

Items
108
Trust grade
B
Status
current

What does a Toolathlon-Verified task look like?

A single task gives the agent one natural-language goal embedded in an already-populated, live multi-application environment — it is a chore to carry out end-to-end via MCP tool calls, not a question to answer in text. Reaching the goal typically takes on the order of 20 sequential tool-calling turns spanning more than one application (e.g. an inbox plus a calendar, or a spreadsheet plus a database). The benchmark's own team has published short summaries of a few illustrative tasks of this shape: one has the agent pull student submissions out of an email inbox and enter the corresponding grades into a course on Canvas (the LMS); another has the agent gather several quarters of institutional-ownership data, adjust the figures for a stock split, and populate a specific spreadsheet template with only the holdings that are common across quarters. In both, correctness is judged purely by whether the live application(s) end up in the correct final state — not by anything the agent says.

How many points on Toolathlon-Verified is real?

On Toolathlon-Verified, a gap smaller than 9.7 points is treated as noise — with 108 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √108 ≈ 9.7, two standard errors on a benchmark this size.

Grade B: No major known issues on Toolathlon-Verified, and the results here hold up — but with 108 items, close calls still deserve an independent check.

7 independent Toolathlon-Verified scores — dots inside one dashed band are closer than Toolathlon-Verified’s 9.7-point noise band, statistically indistinguishable. 10 vendor-reported Toolathlon-Verified scores are not plotted — see the table below.

Can you trust Toolathlon-Verified? Caveats

Toolathlon-Verified is actively discriminating between frontier models — a gap of at least 9.7 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

108 tasks across 32 MCP servers / 600+ tools (per the ICLR 2026 paper, arXiv:2510.25726, the underlying app count is 32 real applications — e.g. Google Calendar, Notion, WooCommerce, Kubernetes, BigQuery — wired to 604 tools total). Each task hands the agent one multi-step chore that needs roughly 20 agent turns on average across several already-populated live applications to finish. Scoring is execution-based, not an LLM judge: every task ships a deterministic per-task verification script that checks whether the final state of the live application(s) matches a ground-truth end state, so the published number is a Pass@1 success rate — the percentage of the 108 tasks whose end-state check passed on a single attempt (the toolathlon.xyz leaderboard also tracks Pass@3 and Pass^3 for repeated-attempt scoring, and reports per-task average turns/tool-calls as a separate efficiency metric). There is no multiple-choice element in this benchmark and therefore no chance floor. Toolathlon is built and maintained by the HKUST NLP group (lead author Junlong Li, advisor Junxian He).

The "Verified" revision hardened the grading and isolated task state after the original Toolathlon shipped some scoring bugs. This is now a named, dated fix rather than a vague claim: Toolathlon-Verified was released 2026-06-30 (per the GitHub repo's release note and the team's own "Introducing Toolathlon-Verified" post on toolathlon.xyz).

What changed between the two versions

That post reports the two versions cover the identical 108 tasks — none added, removed, or renamed — but that 83 of the 108 task packages had net changes across prompts, initial states, ground truth, or evaluators, produced over a repair-and-review span of 339 commits (309 non-merge, 30 merge). It documents specific fixes in both directions: e.g. an evaluator in one task called a logging-lookup helper with positional arguments so a stray parameter shifted into the wrong slot, causing the checker to query the wrong log target and reject otherwise-correct runs; separately, a Python comparison helper in another task returned a (bool, message) tuple that the caller truth-tested directly — since non-empty tuples are always truthy in Python, this let some incorrect answers silently pass.

The team’s own account

Lead author Junlong Li announced the release on X on 2026-07-03, stating the team "corrected tasks, align[ed] graders, isolat[ed] state, harden[ed] the infrastructure, and more — so failures reflect model capabilities" rather than benchmark bugs, while flagging that "Toolathlon-Verified is not perfect." Advisor Junxian He's companion post named the underlying problem the fix addressed: "some live tasks are drifting from the original groundtruth, some tasks are found to have flaws on the design." The revision resets the leaderboard as a new official score series — pre-2026-06-30 Toolathlon scores are not comparable to Toolathlon-Verified numbers (the toolathlon.xyz leaderboard itself states this explicitly).

Reading the leaderboard badges

The toolathlon.xyz leaderboard marks independently-run results with an "evaluated by us" checkmark badge; unbadged numbers are vendor self-reported — the same independent/self-reported split already tracked in this benchmark's score list, e.g. DeepSeek V4 Flash's independently-confirmed +19.8-point jump between its pre-0731 and 0731 minor versions (50.9 to 70.7, both badge-verified on the official board) is evidence the badge distinction is meaningful and not just a labeling formality.

Benchmarks related to Toolathlon-Verified

Is the current lead on Toolathlon-Verified real?

UnverifiedGLM-5.3-Flash leads by 0.6 points, but both scores are vendor-reported — no independent run yet.

Written up in full, for models on this Toolathlon-Verified board: Claude Opus 4.8 vs Gemini 3.1 Pro Preview · Gemini 3.1 Pro Preview vs Kimi K3 — every benchmark each pair shares, not just Toolathlon-Verified.

Toolathlon-Verified leaderboard: current scores

Toolathlon-Verified’s 17 tracked scores were each verified against their source between 2026-06-30 and 2026-09-29.

Toolathlon-Verified scores
ModelToolathlon-Verified score
Kimi K3
Claude Opus 4.8
Muse Spark 1.2
Claude Sonnet 5
DeepSeek V4 Flash (0731)
Gemini 3.1 Pro Preview
GLM-5.2
GLM-5.3-Flash
Claude Sonnet 5.5
MiMo-V2.6-Pro
DeepSeek V4 Pro (0813)
Step 5 Preview
Tencent Hy4 preview
MiMo-V2.6-Flash
Qwen3.8-Flash-Next
GLM-5.3
Qwen3.8-Max

Toolathlon-Verified FAQ

What does Toolathlon-Verified measure?

Realistic multi-step chores that require an agent to juggle many real software tools over a long session, not just answer questions. The benchmark has 108 items.

How big does a gap on Toolathlon-Verified have to be to mean anything?

At least 9.7 points. Across 108 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust Toolathlon-Verified scores?

We grade it B and list it as current — it still tells frontier models apart across 108 items. A gap of at least 9.7 points clears that noise band, but it only counts as a real gap once both scores come from independent runs.

Source: https://github.com/hkust-nlp/Toolathlon