HumanEval.org

Models

Claude Opus 5

Compare with…

HumanEval ratings

Not rated yet — ratings appear with the next official snapshot once this model has eligible votes.

External benchmarks

Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.

General Intelligence

LMArena Text Arena (overall)14881482 14942026-09-02LMArena
LiveBench — Overall80.12026-09-07LiveBench

Reasoning

LiveBench — Reasoning91.22026-09-07LiveBench
GPQA Diamond93.9%91.0% 96.8%2026-07-24Epoch AI Benchmarking Hub

Coding

LiveBench — Coding81.52026-09-07LiveBench

Mathematics

LiveBench — Mathematics95.72026-09-07LiveBench
OTIS Mock AIME 2024–202598.9%96.7% 101.1%2026-07-24Epoch AI Benchmarking Hub

Agentic

LiveBench — Agentic Coding65.22026-09-07LiveBench
Terminal-Bench 4.051.8%48.4% 55.2%2026-07-24Terminal-Bench

Computer use

OSWorld 2.0 (binary accuracy)31.4%2026-09-03OSWorld

Language

LiveBench — Language88.72026-09-07LiveBench

Data analytics

LiveBench — Data Analysis74.52026-09-07LiveBench

Instruction following

LiveBench — Instruction Following63.82026-09-07LiveBench

Benchmark percentile by metric group

Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.

Win rate by opponent

No eligible head-to-head votes yet. Ties count as half a win for each side.

Sample battles

Curated prompts only, with vote tallies — voter identities are never published.

No published sample battles yet — battles appear here once this model has eligible votes on curated prompts.