Toolathlon-Verified
Realistic multi-step chores that require an agent to juggle many real software tools over a long session, not just answer questions.
Items
108
Trust grade
B
Status
current
How many points is real?
On Toolathlon-Verified, a gap smaller than 9.6 points is treated as noise — with 108 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.
Trust caveats
Actively discriminating between frontier models.
108 tasks across 32 MCP servers / 600+ tools. 'Verified' revision hardened the grading and isolated task state after the original Toolathlon shipped some scoring bugs.
Current scores
| Kimi K3 | 76.5independent |
| Claude Opus 4.8 | 76.2independent |
| DeepSeek V4 Pro (0813) | 74.1self-reported |
| Qwen3.8-Max | 72.5self-reported |
| Gemini 3.1 Pro Preview | 61.1independent |