The Model Gap

SWE-bench Verified

Gives an agent a real closed GitHub issue and checks whether its patch fixes it and passes the project's real test suite.

Items
500
Trust grade
D
Status
saturated

How many points is real?

On SWE-bench Verified, a gap smaller than 4.5 points is treated as noise — with 500 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).

In plain English: Grade D: Saturated or known contamination. Rankings here are unreliable regardless of gap size.

Trust caveats

Most frontier models score near the maximum — this benchmark no longer separates them well.

500 tasks, human-filtered from the original SWE-bench with OpenAI. Frontier models now cluster in the mid-90s% — widely regarded as saturated, with independent audits raising contamination concerns at the top of the leaderboard. Treat close rankings here as unreliable; the field is shifting to SWE-bench Pro.

Current scores

Claude Opus 5
DeepSeek V4 Pro (0813)
GPT-5.6 Sol
Kimi K3
Claude Opus 4.8

Source: https://www.swebench.com/verified.html