HumanEval.org

Benchmarks

GPQA Diamond

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelGPQA DiamondObservedSource
GPT-6 AstraOpenAI95.8%93.1% – 98.5%2026-08-30source ↗
Gemini 3.1 Pro PreviewGoogle94.4%91.3% – 97.6%2026-08-06source ↗
Claude Opus 5Anthropic93.9%91.0% – 96.8%2026-07-24source ↗
GPT-5.6 SolOpenAI93.5%90.4% – 96.6%2026-07-09source ↗
GPT-5.6 TerraOpenAI93.3%90.3% – 96.3%2026-07-09source ↗
Grok 4.6SpaceXAI93.2%90.2% – 96.2%2026-08-14source ↗
Kimi K3Moonshot AI93.1%90.2% – 96.0%2026-07-16source ↗
Qwen3.8 Max (0902)Alibaba Qwen92.7%89.4% – 96.0%2026-08-04source ↗
DeepSeek V4 Pro (0813)DeepSeek91.7%88.7% – 94.7%2026-08-18source ↗
GPT-5.6 LunaOpenAI91.6%88.2% – 95.0%2026-07-09source ↗
Claude Opus 4.8Anthropic91.0%87.3% – 94.8%2026-06-07source ↗
GLM-5.3Z.ai90.9%87.8% – 94.0%2026-08-24source ↗
Claude Opus 4.7Anthropic86.4%81.6% – 91.2%2026-08-06source ↗
Claude Fable 5Anthropic85.9%81.0% – 90.7%2026-08-06source ↗
Claude Sonnet 5Anthropic80.3%74.8% – 85.9%2026-08-06source ↗
Llama 4 MaverickMeta67.0%61.5% – 72.5%2025-04-08source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.