HumanEval.org

Benchmarks

OSWorld 2.0 (binary accuracy)

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelOSWorld 2.0 (binary accuracy)ObservedSource
Claude Opus 5Anthropic31.4%2026-09-03source ↗
GPT-5.6 SolOpenAI27.3%2026-09-03source ↗
Claude Opus 4.8Anthropic20.6%2026-09-03source ↗
Claude Opus 4.7Anthropic18.2%2026-09-03source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.