Terminal-Bench 2.1
Whether an AI agent can actually operate a real command line to debug code, administer systems, and fix security issues.
Items
89
Trust grade
B
Status
current
How many points is real?
On Terminal-Bench 2.1, a gap smaller than 10.6 points is treated as noise — with 89 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.
Trust caveats
Actively discriminating between frontier models.
89 tasks; a hardened revision of 2.0 with 26 tasks fixed for bugs and reward-hacking. Small n means single-digit-point gaps are noise.
Current scores
| Kimi K3 | 88.3self-reported |
| Qwen3.8-Max | 86.6self-reported |
| Claude Fable 5 | 83.8independent |
| GLM-5.2 | 81.0self-reported |
| Claude Opus 4.8 | 78.9independent |
Source: https://hub.harborframework.com/datasets/terminal-bench/terminal-bench-2-1/6