HumanEval.org

Models

Claude Opus 4.8not yet evaluated by HumanEval

Compare with…

HumanEval ratings

Not yet evaluated by HumanEval — this model is listed for context from external benchmarks and has not collected blind human votes in the arena. Our rating appears here once it joins the roster and gets eligible votes.

External benchmarks

Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.

General Intelligence

LMArena Text Arena (overall)14821478 14862026-09-02LMArena
LiveBench — Overall76.22026-09-07LiveBench

Reasoning

LiveBench — Reasoning89.22026-09-07LiveBench
GPQA Diamond91.0%87.3% 94.8%2026-06-07Epoch AI Benchmarking Hub

Coding

LiveBench — Coding81.82026-09-07LiveBench

Mathematics

LiveBench — Mathematics94.32026-09-07LiveBench
OTIS Mock AIME 2024–202598.3%95.6% 101.1%2026-06-07Epoch AI Benchmarking Hub

Agentic

LiveBench — Agentic Coding50.52026-09-07LiveBench
Terminal-Bench 4.023.6%20.1% 27.2%2026-05-28Terminal-Bench

Computer use

OSWorld 2.0 (binary accuracy)20.6%2026-09-03OSWorld

Language

LiveBench — Language79.72026-09-07LiveBench

Data analytics

LiveBench — Data Analysis66.02026-09-07LiveBench

Instruction following

LiveBench — Instruction Following72.02026-09-07LiveBench

Benchmark percentile by metric group

Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.