The Model Gap

DeepSWE

Original, never-before-public software engineering tasks, built so agents can't just recall memorized GitHub fixes.

Items
113
Trust grade
B
Status
current

How many points is real?

On DeepSWE, a gap smaller than 9.4 points is treated as noise — with 113 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.

Trust caveats

Actively discriminating between frontier models.

113 tasks across 91 repos, 5 languages. Distinct 2026 benchmark — not the same thing as the earlier 'DeepSWE-Preview' model. Hand-written functional verifiers, low disagreement with independent judges.

Current scores

Claude Opus 5
GPT-5.6 Sol
Claude Fable 5
Kimi K3
DeepSeek V4 Pro (0813)

Source: https://arxiv.org/abs/2607.07946