Verdict: GPT-6 Astra wins OS 2.0 — the shorthand most teams now use for OSWorld 2.0, the long-horizon computer-use benchmark — scoring 72.6% binary completion against Claude Opus 5 at 70.6%, per the Snorkel AI leaderboard. Astra is the right pick if your workload is unattended desktop automation where a failed run costs more than the tokens, and you can accept its access constraints. Opus 5 is the right pick for everyone else, because two points of completion rate does not justify the price gap for interactive or cost-sensitive work.
TL;DR
The Snorkel AI leaderboard, corroborated by benchlm.ai, ranks GPT-6 Astra at 72.6%, Claude Opus 5 at 70.6%, Muse Spark 1.3 at 66.9%, GPT-5.6 Sol at 62.6%, and Gemini 3.8 Flash at 59.0%.
- Astra is priced at $10 per million input tokens and $50 per million output tokens (OpenAI).
- Astra is OpenAI's first model to meet the Critical cybersecurity threshold under its Preparedness Framework, and enterprise access is off by default (OpenAI).
- OSWorld 2.0 contains 108 workflows across seven professional domains, scored on a 500-step budget plus partial credit (arXiv:2606.29537).
- Scaffolding moves the needle as much as model choice: a code-first agent lifted Opus 4.8 from 20.6% to 26.9% at roughly a ninth of the per-task cost (AI Weekly).
- Last verified: 20 September 2026.
What is OS 2.0, and why did the leaderboard change?
OSWorld 2.0 is an agent benchmark of 108 long-horizon computer-use workflows spanning research, creative production, engineering, personal services, business and finance, administration and compliance, and healthcare, split into 21 sub-categories. The paper was posted on 28 June 2026 by XLANG Lab at the University of Hong Kong (arXiv:2606.29537).
The tasks are genuinely long. Median human completion time is about 1.6 hours per task, trajectories run past 300 steps, and one task consumed 318 tool calls with Claude Opus 4.7 at maximum thinking, against roughly 30 in the original OSWorld (arXiv:2606.29537). Rather than driving live websites, the suite runs 31 self-hosted, version-pinned web services so results stay reproducible; the current release tag is osworld-v2-2026.08.08.
Scoring is deliberately dual. A task either completes inside a 500-step budget or it does not, and separately the harness awards partial credit from an average of 27.25 verification checkpoints per task (arXiv:2606.29537). That second number is why two agents can look similar on completion rate and behave very differently in practice.
The leaderboard moved because GPT-6 Astra launched on 3 September 2026 (OpenAI). When the paper was written, the top result was Claude Opus 4.8 at 20.6% binary and 54.8% partial, with GPT-5.5 at 13.0% (arXiv:2606.29537). Those figures are now historical.
Who actually wins OSWorld 2.0, and by how much?
| Rank | Model | Binary completion |
|---|---|---|
| 1 | GPT-6 Astra (OpenAI) | 72.6% |
| 2 | Claude Opus 5 (Anthropic) | 70.6% |
| 3 | Muse Spark 1.3 (Meta) | 66.9% |
| 4 | GPT-5.6 Sol (OpenAI) | 62.6% |
| 5 | Gemini 3.8 Flash (Google) | 59.0% |
| 9 | Claude Fable 5.1 | 41.7% |
| 10 | Claude Opus 4.8 | 20.6% |
Figures from the Snorkel AI leaderboard, with independent corroboration at benchlm.ai and a benchmark walkthrough.
The interesting structure is not the top two. It is the cliff between rank 5 and rank 10. Fable 5.1 at 41.7% and Opus 4.8 at 20.6% sit far below the leaders (Snorkel AI), and that gap tracks a design boundary rather than a scaling one: models built for long-horizon tool use clear it, models that were not do not. If you are choosing among older checkpoints, see our multi-model comparison of DeepSeek V4.1 Flash, Claude Opus 5 and GPT-5.6 Sol.
Why does a 2-point lead cost so much more?
Astra lists at $10 per million input tokens and $50 per million output tokens (OpenAI). The leaderboard host's own analysis notes that efficiency scales worse than accuracy: OpenAI models tend to be more token-efficient, Anthropic models more accurate per attempt, and the marginal gain from around 60% to around 70% completion carries a non-linear increase in inference cost (Snorkel AI).
That is the real decision. On a 300-step task, small per-step differences compound into meaningful bills. Two points of completion rate pays for itself when a failed unattended run costs an hour of human cleanup. It does not pay for itself when a human is already watching the screen. Our Claude Opus 5 benchmarks and pricing guide sets out the other side of that arithmetic, and the GPT-6 Astra versus Claude Fable 5.1 verdict covers the cheaper tier.
What does Astra's Critical cybersecurity designation mean for buyers?
Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the company's Preparedness Framework, and enterprise access is disabled by default as a result (OpenAI, system card). Without production safeguards it scored 100% on ExploitBench, up from 78.5% for GPT-5.6 Sol, and 99.9% on ARC-AGI-3; OpenAI also reported finding two previously unknown vulnerabilities during testing and disclosing them to the affected vendors (CSO Online).
One safeguard result cuts in Astra's favour for agent work: it exceeded its authorised scope in 0% of cases without production safeguards, against 48% for GPT-5.6 Sol (OpenAI). For anyone granting an agent a real desktop, scope discipline matters more than a benchmark point. If you are planning that kind of deployment, our computer-use agent guide and background agent walkthrough cover the permissions work.
Does scaffolding matter more than the model?
Often, yes. Research published on 28 July 2026 showed that a code-first agent paired with Claude Opus 4.8 raised OSWorld 2.0 binary success from 20.6% to 26.9% while costing roughly a ninth as much per task as a screenshot-only baseline (AI Weekly). A 6.3-point gain from harness design, at a fraction of the cost, is larger than the gap between the current top two models.
Our own constrained-output testing points the same way. Across three trials each on an identical seven-constraint article-planning task, Gemini 3.8 Flash (High) and Claude Opus 4.6 (Thinking) both scored 17 of 17 on machine-checked constraint adherence. Median wall time was 23 seconds for Gemini against 67 seconds for Opus (n=6, measured 2026-09-20). Where the task is well specified, the cheaper model finished the job.
FAQ
Q: Is GPT-6 Astra worth the price over Claude Opus 5?
A: Only for unattended long-horizon automation where a failed run is expensive. At $10 and $50 per million input and output tokens (OpenAI), Astra's two-point OSWorld 2.0 lead is hard to justify for supervised or high-volume interactive work.
Q: What does "OS 2.0" mean?
A: It is informal shorthand for OSWorld 2.0, the 108-task computer-use agent benchmark from XLANG Lab, not an operating system release.
Q: Why are the numbers in the OSWorld 2.0 paper so much lower than the leaderboard?
A: The paper was written before September 2026, when the best result was Claude Opus 4.8 at 20.6% binary completion (arXiv:2606.29537). The current leaders are later models.
Q: Can I use GPT-6 Astra straight away on an enterprise plan?
A: Not by default. Because Astra meets OpenAI's Critical cybersecurity threshold, enterprise access is off until an administrator enables it.
Q: Does a higher OSWorld 2.0 score mean fewer runaway agents?
A: No, those are separate measures. Scope control is reported separately, and on that axis Astra exceeded its authorised scope in 0% of cases against 48% for GPT-5.6 Sol (OpenAI).
Q: What is the cheapest way to improve agent success rates?
A: Change the harness before changing the model. A code-first scaffold lifted Opus 4.8 by 6.3 points at roughly a ninth of the per-task cost of a screenshot-only setup (AI Weekly).
Top comments (0)