HumanEval.org

Models

GLM-5.3not yet evaluated by HumanEval

Compare with…

HumanEval ratings

Not yet evaluated by HumanEval — this model is listed for context from external benchmarks and has not collected blind human votes in the arena. Our rating appears here once it joins the roster and gets eligible votes.

External benchmarks

Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.

General Intelligence

LMArena Text Arena (overall)14821475 14892026-09-02LMArena
LiveBench — Overall76.12026-09-07LiveBench

Reasoning

LiveBench — Reasoning85.82026-09-07LiveBench
GPQA Diamond90.9%87.8% 94.0%2026-08-24Epoch AI Benchmarking Hub

Coding

LiveBench — Coding79.02026-09-07LiveBench

Mathematics

LiveBench — Mathematics87.92026-09-07LiveBench
OTIS Mock AIME 2024–202591.1%84.3% 97.9%2026-08-24Epoch AI Benchmarking Hub

Agentic

LiveBench — Agentic Coding60.92026-09-07LiveBench
Terminal-Bench 4.041.8%38.6% 45.0%2026-08-14Terminal-Bench

Language

LiveBench — Language79.92026-09-07LiveBench

Data analytics

LiveBench — Data Analysis70.22026-09-07LiveBench

Instruction following

LiveBench — Instruction Following69.32026-09-07LiveBench

Benchmark percentile by metric group

Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.