HumanEval.org

Models

Llama 4 Maverick

Compare with…

HumanEval ratings

Not rated yet — ratings appear with the next official snapshot once this model has eligible votes.

External benchmarks

Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.

General Intelligence

LMArena Text Arena (overall)13271323 13312026-09-02LMArena

Reasoning

GPQA Diamond67.0%61.5% 72.5%2025-04-08Epoch AI Benchmarking Hub
Humanity's Last Exam5.7%3.9% 7.5%2026-09-08Scale AI SEAL — Humanity's Last Exam

Coding

SWE-bench Verified21.0%2025-07-20SWE-bench

Mathematics

OTIS Mock AIME 2024–202520.6%10.9% 30.3%2025-04-08Epoch AI Benchmarking Hub

Benchmark percentile by metric group

Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.

Win rate by opponent

No eligible head-to-head votes yet. Ties count as half a win for each side.

Sample battles

Curated prompts only, with vote tallies — voter identities are never published.

No published sample battles yet — battles appear here once this model has eligible votes on curated prompts.