OSWorld 2.0

Whether an AI agent can carry out long, multi-app computer workflows on a real Ubuntu desktop and finish them correctly.

Items
108
Trust grade
C
Status
current

What does an OSWorld 2.0 task look like?

One task drops an agent into a fresh Ubuntu virtual machine, preloaded with the files, accounts, and self-hosted mock web apps the scenario needs, and hands it a goal a working professional would recognize: reconcile records scattered across mail, chat, spreadsheets, and a legacy web portal, then produce the correct final artifact. The agent sees the screen only through screenshots and acts only through mouse and keyboard, under a default budget of 500 action steps at 1080p resolution. The suite is long-horizon by design: median skilled-human completion time is about 1.6 hours, and many tasks deliver new information mid-task — a late email that changes the requirements, or a record the agent must recover from an earlier report — so success requires revising plans rather than executing a fixed script. Grading checks a fixed set of weighted checkpoints against the machine's final state, about 27 per task on average, reported two ways: partial reward (mean checkpoint credit earned) and a strict or binary pass (every checkpoint met). No task's actual instructions are reproduced here.

How many points on OSWorld 2.0 is real?

On OSWorld 2.0, a gap smaller than 9.7 points is treated as noise — with 108 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √108 ≈ 9.7, two standard errors on OSWorld 2.0 at this size.

Grade C: Usable, but OSWorld 2.0's 108-item sample is small or the benchmark is nearing saturation — treat close results with extra skepticism.

Can you trust OSWorld 2.0? Caveats

OSWorld 2.0 is actively discriminating between frontier models — a gap of at least 9.7 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.

OSWorld 2.0 is built and maintained by XLANG Lab at the University of Hong Kong with academic and industry collaborators including Snorkel AI; the paper (arXiv:2606.29537) and the first task release (v2026.06.24) appeared 2026-06-26. It ships 108 long-horizon workflows across seven professional domains and 21 sub-categories, grounded in 31 self-hosted mock websites; median skilled-human completion is about 1.6 hours per task. Task classes and assets ship only through gated Hugging Face repositories as an anti-leakage measure; the framework code is Apache-2.0.

Metrics

Every run is graded against weighted checkpoints in the machine's final state (27.25 per task on average). The official board reports two numbers — binary accuracy (every checkpoint met) and partial score (mean checkpoint credit) — under a default 500-step budget; Anthropic's system cards report the same pair as "partial score" and "strict pass rate", pass@1 averaged over five runs. This page pins partial reward as the headline because that is the metric vendors' launch charts quote, and keeps binary/strict in the row notes. The gap is not cosmetic: the same official run (Claude Opus 5, v2.1 full set, max effort) scores 77.67 partial and 44.33 binary.

Version fragmentation is this benchmark's central problem, disclosed here rather than averaged away. The maintainers have published three releases — v2026.06.24 (the paper's set), v2026.08.08 (updated code, task files, assets, and mocked websites), and v2.1 (2026-09-16, a bug-fix release now recommended; status matrix in the repo) — and the official results file carries rows across all three. Vendors report different slices: Anthropic's launch charts say "OSWorld 2.1" (its system cards specify the September 10, 2026 task files) on the combined full set; OpenAI and Google report the offline subset of v2026.08.08 under the official evaluator.

Modified grading

Anthropic's system cards (Fable/Mythos 5.1 §8.14.3, Opus 5.5 §8.13.3, Sonnet 5.5 §8.13.3) disclose that their runs add task-setup and grading-script fixes of their own that they say they reported to the authors, substitute Claude Opus 4.8 wherever a task needs a model grader, and change harness context management (retaining every screenshot, compacting via the Claude API past 100k tokens). The cards state the results are "not directly comparable" to earlier releases or other harness configurations. OpenAI's GPT-6 Astra page answers in a footnote that its Claude figures use the official settings, "not the modified tasks and modified grading from the Fable 5.1 System Card"; Google's Argon methodology omits Anthropic's numbers because they combine the online and offline subsets. OpenAI's quoted Opus 5 offline figure (70.2) matches the authors' own run (70.19) to a decimal — strong evidence the official numbers reproduce.

Setup variance dwarfs small gaps

With 108 items the noise band is roughly ±9.7 points (ceil(100/√108×10)/10, the public methodology formula). Claude Opus 5's official v2.1 full-set partial spans 59.09 to 77.67 on reasoning effort alone; Claude Opus 4.7's original-release partial rises from 20.3 to 49.1 between a 150-step and a 500-step budget.

Trust grade C. In favor: verified runs happen on the maintainers' side (public evaluation requires running the agent with the team and disclosing the implementation); releases pin code, tasks, assets, website, and image versions with task-hash manifests; the results file is dated (updated 2026-10-05) and carries reproduction and trajectory links for its agent entry. Against: the protocol fragmented within three months into irreconcilable vendor configurations, two competitors have put the grading dispute in writing, and grading leans on model graders with each lab substituting its own. No cheating or integrity incident is documented. Until the authors' board absorbs the new models, treat every cross-vendor gap on this page as setup-dependent.

Status (checked 2026-10-08)

The official board's newest tracked rows are Claude Opus 5 and GPT-5.6 Sol; Claude Opus 5.5, Sonnet 5.5, Sonnet 5, Fable 5.1, Fable 5, GPT-6 Astra, GPT-6.1 Sol, and Gemini 4 Argon exist here only as vendor self-reports.

Sources · 11

Is the current lead on OSWorld 2.0 real?

UnverifiedClaude Opus 5.5 leads by 1.1 points, but both scores are vendor-reported — no independent run yet.

Computed on the v2 1 full variant, the same one charted above — other variants of OSWorld 2.0 can rank differently and are shown separately in the table below.

OSWorld 2.0 leaderboard: current scores

OSWorld 2.0’s 17 tracked scores were each verified against their source on 2026-10-08.

OSWorld 2.0 scores — v2026 06 24 full
ModelOSWorld 2.0 score
Claude Opus 4.8
MiniMax M3
OSWorld 2.0 scores — v2026 08 08 full
ModelOSWorld 2.0 score
Claude Opus 5
GPT-5.6 Sol
OSWorld 2.0 scores — v2026 08 08 offline
ModelOSWorld 2.0 score
Claude Opus 5
GPT-5.6 Sol
GPT-6 Astra
GPT-6.1 Sol
Gemini 4 Argon
OSWorld 2.0 scores — v2026 08 08 full modified
ModelOSWorld 2.0 score
Claude Fable 5.1
Claude Opus 5
Claude Fable 5
OSWorld 2.0 scores — v2 1 full
ModelOSWorld 2.0 score
Claude Opus 5
Claude Opus 5.5
Claude Fable 5.1
Claude Sonnet 5.5
Claude Sonnet 5

OSWorld 2.0 FAQ

What does OSWorld 2.0 measure?

Whether an AI agent can carry out long, multi-app computer workflows on a real Ubuntu desktop and finish them correctly. OSWorld 2.0 has 108 items.

How big does a gap on OSWorld 2.0 have to be to mean anything?

At least 9.7 points on OSWorld 2.0. Across OSWorld 2.0's 108 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.

Can you trust OSWorld 2.0 scores?

We grade it C and list it as current — it still tells frontier models apart across 108 OSWorld 2.0 items. A gap of at least 9.7 points clears OSWorld 2.0's own noise band, but it only counts as a real gap once both scores come from independent runs.

What is OSWorld 2.0?

A benchmark of 108 long-horizon computer-use tasks built by XLANG Lab at the University of Hong Kong with collaborators including Snorkel AI. An agent operates a real Ubuntu virtual machine through screenshots plus mouse and keyboard — across files, terminals, and self-hosted mock web apps — and is graded on whether the machine's final state satisfies a weighted set of checkpoints.

Why do OSWorld scores differ between vendors?

Because they are not running the same test. Anthropic reports the v2.1 task files with its own harness changes and its own model as grader; OpenAI and Google report the offline subset of the August 2026 release under the official evaluator. Every score row on this site carries its release, scope, and effort, and numbers from different slices never share a cell.

Is OSWorld 2.1 a different benchmark from OSWorld 2.0?

No. OSWorld 2.1 is a release tag of the same benchmark — a bug-fix refresh the maintainers published on 2026-09-16 and now recommend for all runs. Anthropic markets its launch-chart numbers as "OSWorld 2.1", while the maintainers' own site and paper call the benchmark OSWorld 2.0. The August 2026 release remains active, which is why both labels circulate.

What is the difference between the partial score and the strict score?

Each task is graded against roughly 27 weighted checkpoints in the machine's final state. The partial score is the average share of checkpoint credit earned; the strict pass rate, which the official board calls binary accuracy, counts a task only when every checkpoint is met. The same official run scores 77.67 partial but 44.33 binary, so a partial-reward number always flatters an agent relative to full completion.

How close are the best agents to humans on OSWorld 2.0?

Far off. Skilled humans take a median of about 1.6 hours per task and complete them by construction, while the official board's best binary accuracy is 48.65 percent and the authors report binary completion falls to zero on tasks longer than roughly 163 minutes. Even on partial credit, agents lose track of constraints, miss mid-task updates, and skip verification — the failure modes the benchmark was built to expose.

Which slice of OSWorld does this site count?

All of them, but never mixed. Independent author runs are pinned to a release, a scope (full or offline), and an effort tier, and vendor launch-chart numbers sit in separate rows with their conditions spelled out in the notes. When two rows differ on any of those dimensions, this site treats the comparison as setup-dependent rather than a capability gap.

Can AI models train on OSWorld 2.0 tasks?

The maintainers make that deliberately hard: task classes and assets ship through gated Hugging Face repositories that require accepting an access request, so an evaluated agent cannot easily find task logic online. It is not a rolling-refresh benchmark, though — the 108-task pool has been fixed since June 2026, so a leak that did occur would not self-heal.

Benchmarks related to OSWorld 2.0

Source: https://osworld-v2.xlang.ai/