HumanEval.org

Models

Claude Fable 5not yet evaluated by HumanEval

Compare with…

HumanEval ratings

Not yet evaluated by HumanEval — this model is listed for context from external benchmarks and has not collected blind human votes in the arena. Our rating appears here once it joins the roster and gets eligible votes.

External benchmarks

Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.

General Intelligence

LMArena Text Arena (overall)15071502 15122026-09-02LMArena
LiveBench — Overall83.02026-09-07LiveBench

Reasoning

LiveBench — Reasoning89.72026-09-07LiveBench
GPQA Diamond85.9%81.0% 90.7%2026-08-06Epoch AI Benchmarking Hub

Coding

LiveBench — Coding86.02026-09-07LiveBench

Mathematics

LiveBench — Mathematics96.02026-09-07LiveBench
OTIS Mock AIME 2024–202599.7%99.2% 100.3%2026-06-10Epoch AI Benchmarking Hub

Agentic

LiveBench — Agentic Coding62.22026-09-07LiveBench
Terminal-Bench 4.044.5%40.7% 48.4%2026-06-09Terminal-Bench

Language

LiveBench — Language90.72026-09-07LiveBench

Data analytics

LiveBench — Data Analysis80.52026-09-07LiveBench

Instruction following

LiveBench — Instruction Following75.82026-09-07LiveBench

Benchmark percentile by metric group

Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.