HumanEval.org

Analysis

How our blind human-vote arena ratings relate to external benchmarks, cost, speed and release dates. Our ratings are the primary signal; external numbers are context with provenance. LLMs only — dataset generated 2026-09-23 09:03 UTC.

Analysis appears once official rating snapshots exist. Battles are live in the Arena.

Metric explorer

Plot any two metrics from the board's column universe — our ratings, our battle telemetry, model specs and external benchmarks.

No model has both metrics yet — pick a different pair.

Alibaba QwenAnthropicDeepSeekGoogleMetaMoonshot AIOpenAISpaceXAIZ.ai

0 of 19 models have both metrics. Point size = arena votes received (minimum size = no votes yet); missing data is omitted, never shown as zero.

Dataset: 19 listed LLMs · 0 (model, category) rating pairs · 217 external scores · 11 models with battle telemetry. Telemetry covers all battles including ones excluded from ratings — generation performance is real either way. External scores are the newest observation per model with full provenance on each benchmark page.