HumanEval.org

Models

GPT-5.4not yet evaluated by HumanEval

Compare with…

HumanEval ratings

Not yet evaluated by HumanEval — this model is listed for context from external benchmarks and has not collected blind human votes in the arena. Our rating appears here once it joins the roster and gets eligible votes.

External benchmarks

Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.

General Intelligence

LMArena Text Arena (overall)14771473 14812026-09-02LMArena
LiveBench — Overall78.02026-09-07LiveBench

Reasoning

LiveBench — Reasoning88.12026-09-07LiveBench
Humanity's Last Exam36.2%32.6% 39.9%2026-09-08Scale AI SEAL — Humanity's Last Exam

Coding

LiveBench — Coding77.52026-09-07LiveBench

Mathematics

LiveBench — Mathematics94.22026-09-07LiveBench

Agentic

LiveBench — Agentic Coding53.82026-09-07LiveBench

Language

LiveBench — Language82.62026-09-07LiveBench

Data analytics

LiveBench — Data Analysis79.32026-09-07LiveBench

Instruction following

LiveBench — Instruction Following70.22026-09-07LiveBench

Benchmark percentile by metric group

Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.