SWE-bench Verified
Gives an agent a real closed GitHub issue and checks whether its patch fixes it and passes the project's real test suite.
How many points is real?
On SWE-bench Verified, a gap smaller than 4.5 points is treated as noise — with 500 items, run-to-run variance alone can produce a difference that size. Anything bigger is worth taking seriously (subject to the caveats below).
In plain English: Grade D: Saturated or known contamination. Rankings here are unreliable regardless of gap size.
Trust caveats
Most frontier models score near the maximum — this benchmark no longer separates them well.
500 tasks, human-filtered from the original SWE-bench with OpenAI. Frontier models now cluster in the mid-90s% — widely regarded as saturated, with independent audits raising contamination concerns at the top of the leaderboard. Treat close rankings here as unreliable; the field is shifting to SWE-bench Pro.
Current scores
| Claude Opus 5 | 97.0independent |
| DeepSeek V4 Pro (0813) | 96.4independent |
| GPT-5.6 Sol | 96.2independent |
| Kimi K3 | 93.4independent |
| Claude Opus 4.8 | 88.6independent |