OSWorld 2.0 (binary accuracy)
last retrieved 2026-09-08 21:40 UTC
Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.
| Model | OSWorld 2.0 (binary accuracy) | Observed | Source |
|---|---|---|---|
| Claude Opus 5Anthropic | 31.4% | 2026-09-03 | source ↗ |
| GPT-5.6 SolOpenAI | 27.3% | 2026-09-03 | source ↗ |
| Claude Opus 4.8Anthropic | 20.6% | 2026-09-03 | source ↗ |
| Claude Opus 4.7Anthropic | 18.2% | 2026-09-03 | source ↗ |
— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.
Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.