SWE-bench Verified
Gives an agent a real closed GitHub issue and checks whether its patch fixes it and passes the project's real test suite.
What does a SWE-bench Verified task look like?
One task = a real, previously-closed GitHub issue (bug report or feature request) from one of a fixed set of popular open-source Python repositories, paired with a snapshot of that repo's code at the commit right before the human fix was merged. The AI agent gets the issue description and access to the codebase, has to figure out which files to change and how, and outputs a patch. It never sees the human's actual fix or any later commit history. Grading is automatic: the project's own test suite is run against the patched code, and the task counts as solved only if specific tests that used to fail now pass and nothing that used to pass now breaks.
How many points on SWE-bench Verified is real?
On SWE-bench Verified the 4.5-point noise band is beside the point: the benchmark is saturated, so every comparison across its 500 items is labeled tainted instead of being measured against a threshold.
Grade D: SWE-bench Verified is saturated or has known contamination issues — its rankings are unreliable at any gap size, 500 items or not.
18 independent SWE-bench Verified scores, no noise band drawn: the benchmark is saturated, so every SWE-bench Verified head-to-head is labelled tainted rather than measured against its 4.5-point threshold. How tightly the SWE-bench Verifieddots bunch above is the reason for that call.
Can you trust SWE-bench Verified? Caveats
Most frontier models score near the maximum on SWE-bench Verified — this benchmark no longer separates them well.
500 tasks, human-filtered from the original SWE-bench with OpenAI
SWE-bench Verified was released in August 2024 after 93 professional Python developers screened 1,699 candidate instances from the original SWE-bench pool down to 500, with three independent annotators per sample checking that the issue description is well-specified and that the FAIL_TO_PASS tests don't reject valid solutions (source: OpenAI, "Introducing SWE-bench Verified").
Retired by its own co-creator
Frontier models now cluster in the mid-90s% — widely regarded as saturated, and this is no longer just an outside impression. On 2026-02-23, OpenAI (the benchmark's own co-creator) announced it would stop reporting SWE-bench Verified scores at all, citing three findings: (1) state-of-the-art progress had slowed to a crawl, moving only from 74.9% to 80.9% over the prior six months; (2) an audit of the 138 hardest problems (27.6% of the 500) that OpenAI's own o3 model failed across 64 independent runs found 59.4% of those had flawed test design or underspecified problem statements, making them near-impossible even for a correct solution to pass; and (3) evidence of contamination — frontier models including GPT-5.2, Claude Opus 4.5 and Gemini 3 Flash were able to reproduce gold-standard patches for some tasks near-verbatim, consistent with the public GitHub repos the tasks are drawn from having leaked into training data.
Independent corroboration
This self-audit is distinct from "independent" in the strict sense (OpenAI co-built the benchmark), but genuinely independent work points the same direction: a University of Waterloo study (Prathifkumar, Mathews & Nagappan, arXiv:2512.10218, Dec 2025) found Claude Sonnet 3.5/3.7 located the correct file to edit roughly 6x more often, and resolved issues roughly 3x more often, on SWE-bench Verified than on decontaminated sibling benchmarks (BeetleBox, SWE-rebench) under matched, deliberately minimal-context conditions — a gap best explained by memorization of training data rather than genuine reasoning.
Treat close rankings here as unreliable
Harness changes alone can move a model's own score more than the gap separating adjacent leaderboard ranks (the same harness swap moved claude-sonnet-5's own reported score from a vals.ai baseline of 75.49% to 79.6% in the score data tracked here). The field is shifting to SWE-bench Pro (from Scale AI, cited by OpenAI as its recommended replacement); the gap between the two is large, not small — Scale's own public leaderboard currently tops out around 55-61% for its best-scoring models, well below the mid-90s% top scores tracked here for Verified, underscoring that Pro is a substantially harder, less-contaminated test rather than a like-for-like relabeling of the same skill.
Scoring mechanics
A model gets one attempt to produce a code patch for a real, closed GitHub issue. Grading is fully execution-based — no LLM judge is involved. The patch is applied inside a sandboxed Docker container and the project's real test suite is run, checking both that previously-failing tests now pass (FAIL_TO_PASS) and that previously-passing tests still pass (PASS_TO_PASS, i.e. no regressions introduced). Each of the 500 instances is scored strictly pass/fail with no partial credit; the reported percentage is the "resolved rate" across all 500. There is no multiple-choice structure, so there is no chance floor (a score of 0% is possible). The benchmark is built and maintained by the SWE-bench team at Princeton (Carlos E. Jimenez, John Yang and colleagues; original paper at ICLR 2024, arXiv:2310.06770), with the Verified subset built specifically in collaboration with OpenAI.
Task shape
Each item pairs a real bug report or feature request filed against one of a fixed set of actively-maintained open-source Python repositories with a checkout of that repository pinned to the commit immediately before the human developer's fix was merged. The agent is given the issue text and repository access, must locate and edit the relevant code, and submits a patch; it is never shown the human's actual fix or any git history that postdates the issue. Success is judged purely by whether the submitted patch makes the designated tests pass without breaking any other tests — not by a human or model grader's opinion of code quality or style.
Source status (checked 2026-09-29)
Vals.ai, this column's source, has archived its SWE-bench Verified board as saturated and no longer runs new models on it.
Sources · 8
- OpenAI: "Why SWE-bench Verified no longer measures frontier coding capabilities" (2026-02-23 retirement announcement)
- Pebblous summary of OpenAI's audit (138 problems, 59.4% flawed, 74.9%→80.9%, contamination findings)
- Hacker News discussion thread corroborating the OpenAI announcement
- Prathifkumar, Mathews & Nagappan, "Does SWE-Bench-Verified Test Agent Ability or Model Memory?" (Univ. of Waterloo, arXiv:2512.10218, Dec 2025) — confirms Claude Sonnet 3.5/3.7 as the tested models and the 3x/6x gaps vs BeetleBox/SWE-rebench
- OpenAI: "Introducing SWE-bench Verified" (Aug 2024 methodology: 93 annotators, 1,699 screened to 500, three annotators per sample)
- Jimenez, Yang et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (Princeton, ICLR 2024, arXiv:2310.06770) — original benchmark / maintainer team
- SWE-bench Verified official leaderboard page (task/eval description, maintainer attribution)
- Scale AI, SWE-bench Pro public leaderboard (used to verify the actual Verified-vs-Pro score gap; top scores ~55-61% as of check date, vs mid-90s% on Verified)
Benchmarks related to SWE-bench Verified
- DeepSWE — Long-horizon coding, trust grade B
- LiveCodeBench — Contest coding, trust grade D
Is the current lead on SWE-bench Verified real?
TaintedSWE-bench Verified is saturated — rankings here are unreliable regardless of the gap.
Written up in full, for models on this SWE-bench Verified board: Claude Opus 4.8 vs DeepSeek V4 Pro (0813) · DeepSeek V4 Pro (0813) vs GLM-5.2 · Claude Opus 4.8 vs Gemini 3.1 Pro Preview — every benchmark each pair shares, not just SWE-bench Verified.
SWE-bench Verified scores (not ranked)
SWE-bench Verified’s 18 tracked scores were each verified against their source between 2026-08-17 and 2026-09-03.
| Model | SWE-bench Verified score |
|---|---|
| Claude Opus 5 | 97.0independent |
| DeepSeek V4 Pro (0813) | 96.4independent |
| GPT-5.6 Sol | 96.2independent |
| Grok 4.6 | 95.6independent |
| GLM-5.3 | 95.4independent |
| Claude Fable 5 | 95.0independent |
| Kimi K3 | 93.4independent |
| GPT-5.6 Luna | 93.0independent |
| DeepSeek V4 Flash (0731) | 88.8independent |
| Claude Opus 4.8 | 88.6independent |
| Muse Spark 1.2 | 86.6independent |
| Qwen3.8-Max | 85.6independent |
| GLM-5.2 | 82.8independent |
| Gemini 3.7 Flash | 80.8independent |
| Gemini 3.8 Flash | 80.0independent |
| Claude Sonnet 5 | 79.6independent |
| Gemini 3.1 Pro Preview | 78.8independent |
| MiniMax M3 | 75.0independent |
SWE-bench Verified FAQ
What does SWE-bench Verified measure?
Gives an agent a real closed GitHub issue and checks whether its patch fixes it and passes the project's real test suite. The benchmark has 500 items.
How big does a gap on SWE-bench Verified have to be to mean anything?
No gap on SWE-bench Verified qualifies. Because the benchmark is saturated, we label every head-to-head here tainted and never apply the 4.5-point threshold its 500 items would otherwise imply.
Can you trust SWE-bench Verified scores?
We grade it D and list it as saturated — frontier models now cluster near the maximum, so we label every SWE-bench Verified head-to-head tainted instead of ranking its 500 items.