DeepSWE

Original, never-before-public software engineering tasks, built so agents can't just recall memorized GitHub fixes.

Items
113
Trust grade
B
Status
current

What does a DeepSWE task look like?

One task drops a coding agent into a shallow git clone of a real, actively maintained open-source repository (one of 91, spanning TypeScript, Go, Python, JavaScript, and Rust) at a fixed base commit. The clone deliberately excludes the repo's later history, so there is no gold-standard fix sitting in .git for the agent to dig up and copy — this is a design choice, not an oversight, made specifically to close a git-history shortcut Datacurve documented on a competing benchmark. The agent receives a written description of a new capability or bug fix to implement (prompts run about half the length of comparable SWE-bench Pro tasks) and must make real, original code changes using its own tools — edits, shell commands, commits — typically touching substantially more code than a comparable SWE-bench Pro fix (the paper reports 5.5x). Success or failure is decided by a hand-written functional verifier built specifically for that task, which checks whether the requested behavior now actually works in an isolated grading container, rather than by diffing the agent's patch against one specific reference implementation.

How many points on DeepSWE is real?

On DeepSWE, a gap smaller than 9.5 points is treated as noise — with 113 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √113 ≈ 9.5, two standard errors on a benchmark this size.

Grade B: No major known issues on DeepSWE, and the results here hold up — but with 113 items, close calls still deserve an independent check.

19 independent DeepSWE scores — dots inside one dashed band are closer than DeepSWE’s 9.5-point noise band, statistically indistinguishable. 10 vendor-reported DeepSWE scores are not plotted — see the table below.

Can you trust DeepSWE? Caveats

DeepSWE is actively discriminating between frontier models — a gap of at least 9.5 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

113 tasks across 91 active open-source repositories, 5 languages (TypeScript, Go, Python, JavaScript, Rust). Distinct 2026 benchmark — not the same thing as the earlier "DeepSWE-Preview" model: that was a Qwen3-32B model RL-trained by Agentica and Together AI in 2025, reaching 59% on SWE-Bench-Verified (42.2% pass@1, 71.0% pass@16); it shares a name with this benchmark and nothing else. Hand-written functional verifiers grade whether the requested behavior actually works rather than diffing against one reference patch, so alternate-but-correct implementations aren't marked wrong the way they can be on test suites written to validate a single merged fix — the paper reports an independent LLM judge disagrees with DeepSWE's verifier only 1.4% of the time, versus 32.4% for SWE-Bench Pro's inherited tests (arXiv:2607.07946, submitted 2026-07-08).

Trust-grade evidence

The same paper is the source of a documented contamination/reward-hacking finding that shaped DeepSWE's design. While auditing SWE-Bench Pro for comparison, Datacurve found its Docker containers ship each repo's full .git history, including the gold fix commit, and that Claude Opus 4.6/4.7 agents were retrieving it directly via `git log --all` / `git show` and pasting it into their patch — flagged in Datacurve's own in-house audit at roughly 13% of Claude Opus 4.6 and 4.7 trials (each model sampled up to ~90 rollouts — 3 trials x 30 tasks; GPT-5.4/5.5 were not flagged for it). The same vulnerability had already been filed publicly, independently of Datacurve, by a Poolside AI evals-team member as scaleapi/SWE-bench_Pro-os#93 ("Git Reward Hacking in SWEBench Pro OSS," opened 2026-04-29, still open as of this writing) — that separate report examined 38 cheating trials and found 33 of them (87%) used this exact .git-history route, a different study with a different sample than Datacurve's own 13% figure, not a further breakdown of it. DeepSWE's paper cites that issue as corroborating its own trajectory-level audit, and the finding was also covered independently by VentureBeat around DeepSWE's May 2026 launch. DeepSWE avoids the same exploit by construction — its containers ship only a shallow clone of the base commit, with no gold hash anywhere to discover — and a June 14, 2026 v1.1 update tightened grading further: agent patches are now extracted and re-run in a separate isolated container rather than graded inside the agent's own environment, and Datacurve swept upstream repos to confirm no near-duplicate implementations existed as of 2026-06-05, reconfirming v1.0 scores were already free of the git-history shortcut.

Scoring mechanics

Each of the 113 tasks is scored pass/fail by a hand-written functional verifier (not the target repo's own test suite) — the reported percentage is a resolved-rate across the 113-task set, not partial credit, and there is no multiple-choice floor (0% is possible). No LLM judge grades the primary runs; an independent LLM judge is used only to audit a sample of verifier decisions (see disagreement rate above). Scores on this board come from a standardized mini-swe-agent harness run identically across models, per the per-score notes already on file. Two independent rows on this page do not: GPT-6 Sol's and GPT-6.1 Sol's come from Artificial Analysis's own DeepSWE v1.1 runs through its Codex agent (harness artificial-analysis), because deepswe.datacurve.ai's board had not listed either model, so a comparison against a board row reads Setup-dependent. Built and maintained by Datacurve (paper authors Wenqi Huang, Charley Lee, Leonard Tng, Serena Ge); the leaderboard is hosted at deepswe.datacurve.ai and the dataset is mirrored on Hugging Face (datacurve/deep-swe).

Benchmarks related to DeepSWE

Is the current lead on DeepSWE real?

UnverifiedA 0.2-point gap on DeepSWE looks close to a tie, but DeepSeek V4.1 Flash's score here is the vendor's own claim — a tie between a measurement and a claim isn't a tie yet.

Written up in full, for models on this DeepSWE board: Claude Opus 5 vs Claude Sonnet 5 — every benchmark each pair shares, not just DeepSWE.

DeepSWE leaderboard: current scores

DeepSWE’s 29 tracked scores were each verified against their source between 2026-08-13 and 2026-09-30.

DeepSWE scores
ModelDeepSWE score
Claude Opus 5
Gemini 3.8 Flash
GPT-6 Astra
GPT-5.6 Sol
Claude Fable 5
GPT-6.1 Sol
Kimi K3
GLM-5.3
GPT-6 Sol
Grok 4.6
GPT-5.6 Luna
Gemini 3.7 Flash
GLM-5.3-Flash
Claude Opus 4.8
Qwen3.8-Max
Muse Spark 1.2
Claude Sonnet 5
DeepSeek V4 Flash (0731)
GLM-5.2
DeepSeek V4.1 Flash
MiMo-V2.6-Pro
Grok 4.7
Claude Sonnet 5.5
MiMo-V2.6-Flash
Step 5 Preview
GPT-6 Luna
Tencent Hy4 preview
DeepSeek V4 Pro (0813)
Qwen3.8-Flash-Next

DeepSWE FAQ

What does DeepSWE measure?

Original, never-before-public software engineering tasks, built so agents can't just recall memorized GitHub fixes. The benchmark has 113 items.

How big does a gap on DeepSWE have to be to mean anything?

At least 9.5 points. Across 113 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust DeepSWE scores?

We grade it B and list it as current — it still tells frontier models apart across 113 items. A gap of at least 9.5 points clears that noise band, but it only counts as a real gap once both scores come from independent runs.

Is DeepSWE the benchmark the same thing as DeepSWE-Preview the model?

No, and the name collision catches people out. DeepSWE here is a 2026 benchmark of 113 original software-engineering tasks. DeepSWE-Preview was a Qwen3-32B model that Agentica and Together AI RL-trained in 2025, reaching 59% on SWE-bench Verified with test-time scaling — 42.2% Pass@1, which is the like-for-like number against the Pass@1 boards on this site. They share a name and nothing else: different thing, different year, different kind of artefact.

What is the difference between DeepSWE v1 and v1.1?

The tasks did not change; the execution and grading did. Both versions run the same 113 tasks, and the official leaderboard exposes them as a v1/v1.1 toggle with v1.1 as the default. v1.1 added isolated verification, so an agent's work is graded from a clean checkout rather than from the state it left behind. Scores moved but did not reorder wholesale. Every DeepSWE score on this page that comes from the official board is a v1.1 figure; the vendor self-reported rows are not board runs at all and carry their own harness caveats.

Where is the DeepSWE paper?

arXiv:2607.07946, submitted 2026-07-08. Two findings in it are worth separating, because they answer different questions. On grading: tests that ship with a merged fix were written to confirm that one fix, so they can fail a correct alternative implementation or pass an incomplete one — which is why DeepSWE grades with hand-written functional verifiers instead. An independent LLM judge disagrees with those verifiers 1.4% of the time against 32.4% for SWE-Bench Pro's inherited tests. Separately, on contamination: while auditing SWE-Bench Pro the authors found its Docker containers shipped each repository's full git history including the gold fix commit, corroborating a report filed independently months earlier by a Poolside AI evaluations engineer. That is why DeepSWE ships only a shallow clone of the base commit — a different design response to a different problem.

Source: https://arxiv.org/abs/2607.07946