DeepSWE
Original, never-before-public software engineering tasks, built so agents can't just recall memorized GitHub fixes.
Items
113
Trust grade
B
Status
current
How many points is real?
On DeepSWE, a gap smaller than 9.4 points is treated as noise — with 113 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade B: Reliable, mostly vendor-reported, no major known issues — but still worth an independent check.
Trust caveats
Actively discriminating between frontier models.
113 tasks across 91 repos, 5 languages. Distinct 2026 benchmark — not the same thing as the earlier 'DeepSWE-Preview' model. Hand-written functional verifiers, low disagreement with independent judges.
Current scores
| Claude Opus 5 | 74.0independent |
| GPT-5.6 Sol | 73.0independent |
| Claude Fable 5 | 70.0independent |
| Kimi K3 | 69.0independent |
| DeepSeek V4 Pro (0813) | 62.7self-reported |
Source: https://arxiv.org/abs/2607.07946