HumanEval.org

Benchmarks

LiveBench — Agentic Coding

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelLiveBench — Agentic CodingObservedSource
Claude Fable 5.1Anthropic66.12026-09-07source ↗
Claude Opus 5Anthropic65.22026-09-07source ↗
Qwen3.8 Max (0902)Alibaba Qwen64.72026-09-07source ↗
Claude Fable 5Anthropic62.22026-09-07source ↗
Kimi K3Moonshot AI62.22026-09-07source ↗
GLM-5.3Z.ai60.92026-09-07source ↗
Claude Sonnet 5Anthropic59.42026-09-07source ↗
GPT-6 AstraOpenAI57.32026-09-07source ↗
Grok 4.6SpaceXAI57.02026-09-07source ↗
GPT-5.6 SolOpenAI56.22026-09-07source ↗
DeepSeek V4 Pro (0813)DeepSeek55.02026-09-07source ↗
GPT-5.6 TerraOpenAI55.02026-09-07source ↗
Gemini 3.8 FlashGoogle54.22026-09-07source ↗
GPT-5.4OpenAI53.82026-09-07source ↗
Claude Opus 4.7Anthropic50.72026-09-07source ↗
Claude Opus 4.8Anthropic50.52026-09-07source ↗
GPT-5.6 LunaOpenAI48.42026-09-07source ↗
Gemini 3.1 Pro PreviewGoogle44.12026-09-07source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: score, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.