OSWorld 2.0
Whether an AI agent can carry out long, multi-app computer workflows on a real Ubuntu desktop and finish them correctly.
What does an OSWorld 2.0 task look like?
One task drops an agent into a fresh Ubuntu virtual machine, preloaded with the files, accounts, and self-hosted mock web apps the scenario needs, and hands it a goal a working professional would recognize: reconcile records scattered across mail, chat, spreadsheets, and a legacy web portal, then produce the correct final artifact. The agent sees the screen only through screenshots and acts only through mouse and keyboard, under a default budget of 500 action steps at 1080p resolution. The suite is long-horizon by design: median skilled-human completion time is about 1.6 hours, and many tasks deliver new information mid-task — a late email that changes the requirements, or a record the agent must recover from an earlier report — so success requires revising plans rather than executing a fixed script. Grading checks a fixed set of weighted checkpoints against the machine's final state, about 27 per task on average, reported two ways: partial reward (mean checkpoint credit earned) and a strict or binary pass (every checkpoint met). No task's actual instructions are reproduced here.
How many points on OSWorld 2.0 is real?
On OSWorld 2.0, a gap smaller than 9.7 points is treated as noise — with 108 items, run-to-run variance alone can produce a difference that size. The floor itself is derived, not asserted: 100 ÷ √108 ≈ 9.7, two standard errors on OSWorld 2.0 at this size.
Grade C: Usable, but OSWorld 2.0's 108-item sample is small or the benchmark is nearing saturation — treat close results with extra skepticism.
Can you trust OSWorld 2.0? Caveats
OSWorld 2.0 is actively discriminating between frontier models — a gap of at least 9.7 points here is worth taking seriously, though it only counts as a confirmed real gap once both scores come from independent runs.
OSWorld 2.0 is built and maintained by XLANG Lab at the University of Hong Kong with academic and industry collaborators including Snorkel AI; the paper (arXiv:2606.29537) and the first task release (v2026.06.24) appeared 2026-06-26. It ships 108 long-horizon workflows across seven professional domains and 21 sub-categories, grounded in 31 self-hosted mock websites; median skilled-human completion is about 1.6 hours per task. Task classes and assets ship only through gated Hugging Face repositories as an anti-leakage measure; the framework code is Apache-2.0.
Metrics
Every run is graded against weighted checkpoints in the machine's final state (27.25 per task on average). The official board reports two numbers — binary accuracy (every checkpoint met) and partial score (mean checkpoint credit) — under a default 500-step budget; Anthropic's system cards report the same pair as "partial score" and "strict pass rate", pass@1 averaged over five runs. This page pins partial reward as the headline because that is the metric vendors' launch charts quote, and keeps binary/strict in the row notes. The gap is not cosmetic: the same official run (Claude Opus 5, v2.1 full set, max effort) scores 77.67 partial and 44.33 binary.
Version fragmentation is this benchmark's central problem, disclosed here rather than averaged away. The maintainers have published three releases — v2026.06.24 (the paper's set), v2026.08.08 (updated code, task files, assets, and mocked websites), and v2.1 (2026-09-16, a bug-fix release now recommended; status matrix in the repo) — and the official results file carries rows across all three. Vendors report different slices: Anthropic's launch charts say "OSWorld 2.1" (its system cards specify the September 10, 2026 task files) on the combined full set; OpenAI and Google report the offline subset of v2026.08.08 under the official evaluator.
Modified grading
Anthropic's system cards (Fable/Mythos 5.1 §8.14.3, Opus 5.5 §8.13.3, Sonnet 5.5 §8.13.3) disclose that their runs add task-setup and grading-script fixes of their own that they say they reported to the authors, substitute Claude Opus 4.8 wherever a task needs a model grader, and change harness context management (retaining every screenshot, compacting via the Claude API past 100k tokens). The cards state the results are "not directly comparable" to earlier releases or other harness configurations. OpenAI's GPT-6 Astra page answers in a footnote that its Claude figures use the official settings, "not the modified tasks and modified grading from the Fable 5.1 System Card"; Google's Argon methodology omits Anthropic's numbers because they combine the online and offline subsets. OpenAI's quoted Opus 5 offline figure (70.2) matches the authors' own run (70.19) to a decimal — strong evidence the official numbers reproduce.
Setup variance dwarfs small gaps
With 108 items the noise band is roughly ±9.7 points (ceil(100/√108×10)/10, the public methodology formula). Claude Opus 5's official v2.1 full-set partial spans 59.09 to 77.67 on reasoning effort alone; Claude Opus 4.7's original-release partial rises from 20.3 to 49.1 between a 150-step and a 500-step budget.
Trust grade C. In favor: verified runs happen on the maintainers' side (public evaluation requires running the agent with the team and disclosing the implementation); releases pin code, tasks, assets, website, and image versions with task-hash manifests; the results file is dated (updated 2026-10-05) and carries reproduction and trajectory links for its agent entry. Against: the protocol fragmented within three months into irreconcilable vendor configurations, two competitors have put the grading dispute in writing, and grading leans on model graders with each lab substituting its own. No cheating or integrity incident is documented. Until the authors' board absorbs the new models, treat every cross-vendor gap on this page as setup-dependent.
Status (checked 2026-10-08)
The official board's newest tracked rows are Claude Opus 5 and GPT-5.6 Sol; Claude Opus 5.5, Sonnet 5.5, Sonnet 5, Fable 5.1, Fable 5, GPT-6 Astra, GPT-6.1 Sol, and Gemini 4 Argon exist here only as vendor self-reports.
Sources · 11
- XLANG Lab (HKU) — OSWorld 2.0 project site and verified leaderboard
- Official leaderboard results file (51 rows across three releases; updated 2026-10-05; checked 2026-10-08)
- Yuan et al. — OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks (arXiv:2606.29537, June 2026)
- GitHub — xlang-ai/OSWorld-V2 (release log: 2026-06-24, 2026-08-08, v2.1 on 2026-09-16; verified-leaderboard process; gated task distribution)
- Anthropic — Claude Opus 5.5 launch (2026-09-22): OSWorld 2.1 chart, 81.8 / 80.7 / 74.0 partial
- Anthropic — Claude Opus 5.5 System Card §8.13.3 (Sep 10 2026 task files, retained-screenshot context management, Opus 4.8 model grader; strict pass rates)
- Anthropic — Claude Sonnet 5.5 launch and System Card §8.13.3 (2026-09-28): 80.1 / 57.0 partial, 43.5 / 25.6 strict, same configuration as Opus 5.5 card
- Anthropic — Fable/Mythos 5.1 launch and System Card §8.14.3 (August 2026 task release plus Anthropic's own task and grading fixes; 77.9 / 72.9 / 75.4 partial, 41.7 / 36.1 / 39.6 strict)
- OpenAI — Introducing GPT-6 Astra (2026-09-03): OSWorld 2.0 offline-set partial table; footnote states Claude figures use official settings, not the Fable 5.1 System Card's modified tasks and grading
- OpenAI — Introducing GPT-6.1 Sol (2026-09-29): chart caption "partial reward on the offline set from the v2026.08.08 release"; Sol max 71.42, Astra max 73.49
- Google DeepMind — Gemini models page and Gemini 4 Argon evals methodology (2026-09-30): offline-subset partial 69.2, self-computed max-over-3-runs, official evaluator with the team's 08.08 patch
Is the current lead on OSWorld 2.0 real?
UnverifiedClaude Opus 5.5 leads by 1.1 points, but both scores are vendor-reported — no independent run yet.
Computed on the v2 1 full variant, the same one charted above — other variants of OSWorld 2.0 can rank differently and are shown separately in the table below.
OSWorld 2.0 leaderboard: current scores
OSWorld 2.0’s 17 tracked scores were each verified against their source on 2026-10-08.
| Model | OSWorld 2.0 score |
|---|---|
| Claude Opus 4.8 | 54.8independent |
| MiniMax M3 | 22.3independent |
| Model | OSWorld 2.0 score |
|---|---|
| Claude Opus 5 | 68.3independent |
| GPT-5.6 Sol | 62.7independent |
| Model | OSWorld 2.0 score |
|---|---|
| Claude Opus 5 | 70.2independent |
| GPT-5.6 Sol | 64.1independent |
| GPT-6 Astra | 72.6self-reported |
| GPT-6.1 Sol | 71.4self-reported |
| Gemini 4 Argon | 69.2self-reported |
| Model | OSWorld 2.0 score |
|---|---|
| Claude Fable 5.1 | 77.9self-reported |
| Claude Opus 5 | 75.4self-reported |
| Claude Fable 5 | 72.9self-reported |
| Model | OSWorld 2.0 score |
|---|---|
| Claude Opus 5 | 77.7independent |
| Claude Opus 5.5 | 81.8self-reported |
| Claude Fable 5.1 | 80.7self-reported |
| Claude Sonnet 5.5 | 80.1self-reported |
| Claude Sonnet 5 | 57.0self-reported |
OSWorld 2.0 FAQ
What does OSWorld 2.0 measure?
Whether an AI agent can carry out long, multi-app computer workflows on a real Ubuntu desktop and finish them correctly. OSWorld 2.0 has 108 items.
How big does a gap on OSWorld 2.0 have to be to mean anything?
At least 9.7 points on OSWorld 2.0. Across OSWorld 2.0's 108 items, a smaller gap sits inside the spread that repeat runs produce on their own, so we call it a tie — provided both scores are independent runs; if either side is a vendor claim, the row is labelled unverified instead.
Can you trust OSWorld 2.0 scores?
We grade it C and list it as current — it still tells frontier models apart across 108 OSWorld 2.0 items. A gap of at least 9.7 points clears OSWorld 2.0's own noise band, but it only counts as a real gap once both scores come from independent runs.
What is OSWorld 2.0?
A benchmark of 108 long-horizon computer-use tasks built by XLANG Lab at the University of Hong Kong with collaborators including Snorkel AI. An agent operates a real Ubuntu virtual machine through screenshots plus mouse and keyboard — across files, terminals, and self-hosted mock web apps — and is graded on whether the machine's final state satisfies a weighted set of checkpoints.
Why do OSWorld scores differ between vendors?
Because they are not running the same test. Anthropic reports the v2.1 task files with its own harness changes and its own model as grader; OpenAI and Google report the offline subset of the August 2026 release under the official evaluator. Every score row on this site carries its release, scope, and effort, and numbers from different slices never share a cell.
Is OSWorld 2.1 a different benchmark from OSWorld 2.0?
No. OSWorld 2.1 is a release tag of the same benchmark — a bug-fix refresh the maintainers published on 2026-09-16 and now recommend for all runs. Anthropic markets its launch-chart numbers as "OSWorld 2.1", while the maintainers' own site and paper call the benchmark OSWorld 2.0. The August 2026 release remains active, which is why both labels circulate.
What is the difference between the partial score and the strict score?
Each task is graded against roughly 27 weighted checkpoints in the machine's final state. The partial score is the average share of checkpoint credit earned; the strict pass rate, which the official board calls binary accuracy, counts a task only when every checkpoint is met. The same official run scores 77.67 partial but 44.33 binary, so a partial-reward number always flatters an agent relative to full completion.
How close are the best agents to humans on OSWorld 2.0?
Far off. Skilled humans take a median of about 1.6 hours per task and complete them by construction, while the official board's best binary accuracy is 48.65 percent and the authors report binary completion falls to zero on tasks longer than roughly 163 minutes. Even on partial credit, agents lose track of constraints, miss mid-task updates, and skip verification — the failure modes the benchmark was built to expose.
Which slice of OSWorld does this site count?
All of them, but never mixed. Independent author runs are pinned to a release, a scope (full or offline), and an effort tier, and vendor launch-chart numbers sit in separate rows with their conditions spelled out in the notes. When two rows differ on any of those dimensions, this site treats the comparison as setup-dependent rather than a capability gap.
Can AI models train on OSWorld 2.0 tasks?
The maintainers make that deliberately hard: task classes and assets ship through gated Hugging Face repositories that require accepting an access request, so an evaluated agent cannot easily find task logic online. It is not a rolling-refresh benchmark, though — the 108-task pool has been fixed since June 2026, so a leak that did occur would not self-heal.
Benchmarks related to OSWorld 2.0
- Agents' Last Exam — Professional work, trust grade A
- Toolathlon-Verified — Multi-tool chores, trust grade B
Source: https://osworld-v2.xlang.ai/