HumanEval.org

Benchmarks

LMArena Text Arena (overall)

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelLMArena Text Arena (overall)ObservedSource
Claude Fable 5Anthropic15071502 – 15122026-09-02source ↗
Claude Fable 5.1Anthropic15041493 – 15152026-09-02source ↗
Claude Opus 4.7Anthropic15021498 – 15062026-09-02source ↗
Gemini 3.8 FlashGoogle14941485 – 15032026-09-02source ↗
Kimi K3Moonshot AI14891484 – 14942026-09-02source ↗
Claude Opus 5Anthropic14881482 – 14942026-09-02source ↗
Gemini 3.1 Pro PreviewGoogle14871484 – 14902026-09-02source ↗
GPT-5.6 SolOpenAI14831478 – 14882026-09-02source ↗
Claude Opus 4.8Anthropic14821478 – 14862026-09-02source ↗
GLM-5.3Z.ai14821475 – 14892026-09-02source ↗
Qwen3.8 Max (0902)Alibaba Qwen14801474 – 14862026-09-02source ↗
GPT-5.4OpenAI14771473 – 14812026-09-02source ↗
GPT-5.6 TerraOpenAI14661461 – 14712026-09-02source ↗
Claude Sonnet 5Anthropic14621457 – 14672026-09-02source ↗
Grok 4.6SpaceXAI14611451 – 14712026-09-02source ↗
DeepSeek V4 Pro (0813)DeepSeek14601452 – 14682026-09-02source ↗
GPT-5.6 LunaOpenAI14531448 – 14582026-09-02source ↗
Llama 4 MaverickMeta13271323 – 13312026-09-02source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: Elo, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.