HumanEval.org

Benchmarks

Humanity's Last Exam

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelHumanity's Last ExamObservedSource
Claude Fable 5.1Anthropic46.5%42.6% – 50.4%2026-09-08source ↗
Gemini 3.1 Pro PreviewGoogle46.4%42.6% – 50.3%2026-09-08source ↗
GPT-5.4OpenAI36.2%32.6% – 39.9%2026-09-08source ↗
Claude Opus 4.7Anthropic36.2%32.5% – 39.9%2026-09-08source ↗
Llama 4 MaverickMeta5.7%3.9% – 7.5%2026-09-08source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.