The Model Gap

Toolathlon-Verified

Realistic multi-step chores that require an agent to juggle many real software tools over a long session, not just answer questions.

Items
108
Trust grade
B
Status
current

How many points is real?

On Toolathlon-Verified, a gap smaller than 9.6 points is treated as noise — with 108 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.

Trust caveats

Actively discriminating between frontier models.

108 tasks across 32 MCP servers / 600+ tools. 'Verified' revision hardened the grading and isolated task state after the original Toolathlon shipped some scoring bugs.

Current scores

Kimi K3
Claude Opus 4.8
DeepSeek V4 Pro (0813)
Qwen3.8-Max
Gemini 3.1 Pro Preview

Source: https://github.com/hkust-nlp/Toolathlon