Terminal-Bench 2.1

Whether an AI agent can actually operate a real command line to debug code, administer systems, and fix security issues.

Items
89
Trust grade
B
Status
current

What does a Terminal-Bench 2.1 task look like?

One task instantiates a fresh, isolated container/VM environment and gives the agent a single natural-language objective — for example, get a broken build compiling, close off a mis-configured service's known vulnerability, or stand up and wire together a small data-processing or model-training pipeline — with no further hand-holding. The agent works purely through a terminal: it can run arbitrary multi-turn shell commands, inspect output, and iterate, the same way a human sysadmin or engineer would. When it finishes (or times out), grading runs a hidden, task-specific automated test suite against the resulting state of the environment — files written, services running, data produced — rather than against anything the agent said; there is no reference answer to text-match and no credit for output that merely looks right. This describes the general task format the maintainers themselves describe (compiling code, configuring servers, training models, debugging systems, fixing security issues) — not any single specific item's actual instructions.

How many points on Terminal-Bench 2.1 is real?

On Terminal-Bench 2.1, a gap smaller than 10.6 points is treated as noise — with 89 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √89 ≈ 10.6, two standard errors on a benchmark this size.

Grade B: No major known issues on Terminal-Bench 2.1, and the results here hold up — but with 89 items, close calls still deserve an independent check.

Can you trust Terminal-Bench 2.1? Caveats

Terminal-Bench 2.1 is actively discriminating between frontier models — a gap of at least 10.6 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

89 tasks; a hardened revision of 2.0 with 26 tasks fixed for bugs and reward-hacking (many of the fixes were carried over directly from Z.ai's own Terminal-Bench 2.0 Verified patch set). Small n means single-digit-point gaps are noise — with only 89 items, treat anything inside roughly a 10-point band as statistically indistinguishable.

Trust grade B rests on a documented, dated integrity process rather than a clean record. tbench.ai's own "Leaderboard Integrity Update" (19 April 2026) disclosed that two submissions were pulled for cheating: OpenBlock's "OB-1" for storing encrypted solutions inside its agent binary, and QuantFlow's "Pilot" for uploading the tasks' own test files as part of its agent setup (both incidents distinct from OB-1's earlier, separately-flagged timeout-modification issue on Terminal-Bench 1.0). A third submission, ForgeCode, had its passing trials rescored to zero for reward-hacking after its agent was found pulling a task's solution from the internet and pasting it into its own agent-instructions file. These violations were surfaced by outside auditors (named as Adam Stein, Davis Brown, and the Ante team), not caught internally by the maintainers first, and the response was to require public trajectories for every passing trial going forward, add an agent-judge review layer, and zero out (not merely flag) any hacked trial. That is real evidence both that agents actively try to game this benchmark and that there is a named, functioning correction process — the reason this sits at B rather than lower.

Scoring mechanics

Each of the 89 tasks drops an agent into an isolated sandboxed environment with a natural-language goal and terminal access; a task counts as passed only if the agent's final state clears every test in that task's hidden, task-specific automated test suite (pytest-based) — there is no LLM judge grading transcripts, no partial credit for a correct-looking-but-failing state, and no chance floor since nothing is multiple choice. Official leaderboard submissions run multiple trials per task (the published rules require at least 5), and results are reported as pass@1 accuracy (percent of trials passed); vendor self-reports can use other sampling regimes, such as avg@10 (averaged over 10 trials), which is not the same as a single official pass@1 run. Terminal-Bench is a joint project of Stanford University and the Laude Institute, publicly introduced 19 May 2025, and is now hosted through the open-source Harbor framework at tbench.ai; its reference evaluation harness is Terminus (now on its second major version, "Terminus 2"), a neutral agent scaffold released alongside the benchmark so any model can be run under one shared methodology.

Harness disclosure

The score list above mixes four different measurement conditions under one column — Artificial Analysis's own harness runs, tbench.ai's official leaderboard-verified runs (agent scaffold varies by submission, e.g. Codex for the top-ranked GPT-6 Astra entry and Claude Code for Fable 5's), twelve vals.ai runs (Terminus 2, from its archived table), and one unverified vendor self-report (Step 5 Preview, harness not stated) — none of this is directly comparable on harness grounds alone, and how to formally separate or label these conditions in the schema is a pending decision, not resolved by this note.

Source status (checked 2026-09-30)

Only vals.ai has announced a stop. It has archived its Terminal-Bench 2.1 board, stating that Terminal-Bench 4.0 replaces it in the Coding sector of its Vals Index and that it no longer runs new model releases on it; Claude Sonnet 5.5 is the newest model in the archived table. The other two independent sources have added nothing recent: tbench.ai's own 2.1 leaderboard lists 18 entries, the newest GPT-6 Astra (released 2026-09-03), and Artificial Analysis has published no Terminal-Bench 2.1 value for any model released after 2026-09-11, GPT-6.1 Sol included, while showing a Terminal-Bench 4.0 value for GPT-6.1 Sol. Unless one of them resumes, a newly released model's cell in this column can only hold a vendor self-report or stay empty.

vals.ai as a source (added 2026-10-01): its archived table now supplies twelve rows here: Claude Sonnet 5.5, added 2026-09-29, and eleven more on 2026-10-01. Four fill cells that were empty (Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, Grok 4.7) and seven replace vendor self-reports that ran 7.4 to 30.3 points higher: Kimi K3 88.3 to 80.90, GLM-5.2 81.0 to 67.79, MiniMax M3 66 to 53.56, MiMo-V2.6-Flash 87.6 to 76.40, DeepSeek V4.1 Flash 90.6 to 74.53, MiMo-V2.6-Pro 89.9 to 67.79, and Hy4 preview 85.4 to 55.06. Where a model already had an independent row from Artificial Analysis or tbench.ai, that row stays and vals.ai's figure goes in its row notes; the two can differ by up to 24 points (DeepSeek V4 Pro 0813: 78.7 on AA, 54.68 on vals.ai). Effort differs by row (Claude models at high, the OpenAI models and GLM-5.2 at max, Grok 4.7 at xhigh, DeepSeek V4.1 Flash at high; Kimi, MiniMax, MiMo and Hy4 with no effort parameter listed), and each row names its setting. vals.ai's DeepSeek V4 Flash 0731 row (67.04) is excluded: it ran with thinking off, unlike every other row here. Claude rows carry vals.ai's fallback caveat: Opus 5.5 used Opus 5 and Opus 4.8 on 26 of 267 tasks, and counting those as failures lowers it from 87.64 to 79.77.

Benchmarks related to Terminal-Bench 2.1

Is the current lead on Terminal-Bench 2.1 real?

TieA 2.3-point gap on Terminal-Bench 2.1 (n=89) is inside the noise band (±10.6) — call it a tie.

Written up in full, for models on this Terminal-Bench 2.1 board: Claude Opus 5 vs Claude Sonnet 5 · Gemini 3.1 Pro Preview vs Kimi K3 · Claude Opus 4.8 vs DeepSeek V4 Pro (0813) — every benchmark each pair shares, not just Terminal-Bench 2.1.

Terminal-Bench 2.1 leaderboard: current scores

Terminal-Bench 2.1’s 32 tracked scores were each verified against their source between 2026-08-17 and 2026-10-01.

Terminal-Bench 2.1 scores
ModelTerminal-Bench 2.1 score
Claude Fable 5.1
Claude Opus 5
Qwen3.8-Max
Grok 4.6
GPT-5.6 Sol
Claude Opus 5.5
Gemini 3.8 Flash
GPT-6 Astra
Qwen3.8-Flash-Next
Gemini 3.7 Flash
GLM-5.3-Flash
Muse Spark 1.3
GLM-5.3
Claude Fable 5
Claude Sonnet 5.5
GPT-6 Sol
Kimi K3
GPT-5.6 Luna
Muse Spark 1.2
Claude Opus 4.8
DeepSeek V4 Pro (0813)
MiMo-V2.6-Flash
Claude Sonnet 5
DeepSeek V4.1 Flash
Grok 4.7
GPT-6 Luna
GLM-5.2
MiMo-V2.6-Pro
Gemini 3.1 Pro Preview
Tencent Hy4 preview
MiniMax M3
Step 5 Preview

Terminal-Bench 2.1 FAQ

What does Terminal-Bench 2.1 measure?

Whether an AI agent can actually operate a real command line to debug code, administer systems, and fix security issues. The benchmark has 89 items.

How big does a gap on Terminal-Bench 2.1 have to be to mean anything?

At least 10.6 points. Across 89 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust Terminal-Bench 2.1 scores?

We grade it B and list it as current — it still tells frontier models apart across 89 items. A gap of at least 10.6 points clears that noise band, but it only counts as a real gap once both scores come from independent runs.

Source: https://hub.harborframework.com/datasets/terminal-bench/terminal-bench-2-1/6