HumanEval.org

Models

Gemini 3.1 Pro Previewnot yet evaluated by HumanEval

Compare with…

HumanEval ratings

Not yet evaluated by HumanEval — this model is listed for context from external benchmarks and has not collected blind human votes in the arena. Our rating appears here once it joins the roster and gets eligible votes.

External benchmarks

Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.

General Intelligence

LMArena Text Arena (overall)14871484 14902026-09-02LMArena
LiveBench — Overall77.02026-09-07LiveBench

Reasoning

LiveBench — Reasoning84.02026-09-07LiveBench
GPQA Diamond94.4%91.3% 97.6%2026-08-06Epoch AI Benchmarking Hub
Humanity's Last Exam46.4%42.6% 50.3%2026-09-08Scale AI SEAL — Humanity's Last Exam

Coding

LiveBench — Coding76.52026-09-07LiveBench

Mathematics

LiveBench — Mathematics91.02026-09-07LiveBench
OTIS Mock AIME 2024–202595.6%89.5% 101.6%2026-08-06Epoch AI Benchmarking Hub

Agentic

LiveBench — Agentic Coding44.12026-09-07LiveBench

Language

LiveBench — Language85.42026-09-07LiveBench

Data analytics

LiveBench — Data Analysis78.52026-09-07LiveBench

Instruction following

LiveBench — Instruction Following79.12026-09-07LiveBench

Benchmark percentile by metric group

Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.