The Model Gap

Terminal-Bench 2.1

Whether an AI agent can actually operate a real command line to debug code, administer systems, and fix security issues.

Items
89
Trust grade
B
Status
current

How many points is real?

On Terminal-Bench 2.1, a gap smaller than 10.6 points is treated as noise — with 89 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.

Trust caveats

Actively discriminating between frontier models.

89 tasks; a hardened revision of 2.0 with 26 tasks fixed for bugs and reward-hacking. Small n means single-digit-point gaps are noise.

Current scores

Kimi K3
Qwen3.8-Max
Claude Fable 5
GLM-5.2
Claude Opus 4.8

Source: https://hub.harborframework.com/datasets/terminal-bench/terminal-bench-2-1/6