Claude Opus 5
HumanEval ratings
Not rated yet — ratings appear with the next official snapshot once this model has eligible votes.
External benchmarks
Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.
General Intelligence
Reasoning
| LiveBench — Reasoning | 91.2 | 2026-09-07 | LiveBench |
| GPQA Diamond | 93.9%91.0% – 96.8% | 2026-07-24 | Epoch AI Benchmarking Hub |
Coding
| LiveBench — Coding | 81.5 | 2026-09-07 | LiveBench |
Mathematics
| LiveBench — Mathematics | 95.7 | 2026-09-07 | LiveBench |
| OTIS Mock AIME 2024–2025 | 98.9%96.7% – 101.1% | 2026-07-24 | Epoch AI Benchmarking Hub |
Agentic
| LiveBench — Agentic Coding | 65.2 | 2026-09-07 | LiveBench |
| Terminal-Bench 4.0 | 51.8%48.4% – 55.2% | 2026-07-24 | Terminal-Bench |
Computer use
| OSWorld 2.0 (binary accuracy) | 31.4% | 2026-09-03 | OSWorld |
Language
| LiveBench — Language | 88.7 | 2026-09-07 | LiveBench |
Data analytics
| LiveBench — Data Analysis | 74.5 | 2026-09-07 | LiveBench |
Instruction following
| LiveBench — Instruction Following | 63.8 | 2026-09-07 | LiveBench |
Benchmark percentile by metric group
Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.
Win rate by opponent
No eligible head-to-head votes yet. Ties count as half a win for each side.
Sample battles
Curated prompts only, with vote tallies — voter identities are never published.
No published sample battles yet — battles appear here once this model has eligible votes on curated prompts.