HumanEval.org

Benchmarks

OTIS Mock AIME 2024–2025

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelOTIS Mock AIME 2024–2025ObservedSource
Claude Fable 5.1Anthropic100.0%100.0% – 100.0%2026-09-01source ↗
GPT-5.6 SolOpenAI100.0%100.0% – 100.0%2026-07-09source ↗
GPT-6 AstraOpenAI100.0%100.0% – 100.0%2026-08-30source ↗
Claude Fable 5Anthropic99.7%99.2% – 100.3%2026-06-10source ↗
GPT-5.6 TerraOpenAI99.7%99.2% – 100.3%2026-07-09source ↗
Qwen3.8 Max (0902)Alibaba Qwen99.4%98.7% – 100.2%2026-08-04source ↗
Grok 4.6SpaceXAI99.2%98.3% – 100.1%2026-08-14source ↗
Claude Opus 5Anthropic98.9%96.7% – 101.1%2026-07-24source ↗
DeepSeek V4 Pro (0813)DeepSeek98.6%96.7% – 100.5%2026-08-18source ↗
Claude Opus 4.8Anthropic98.3%95.6% – 101.1%2026-06-07source ↗
GPT-5.6 LunaOpenAI98.3%96.0% – 100.6%2026-07-09source ↗
Kimi K3Moonshot AI97.2%95.0% – 99.4%2026-07-16source ↗
Gemini 3.1 Pro PreviewGoogle95.6%89.5% – 101.6%2026-08-06source ↗
GLM-5.3Z.ai91.1%84.3% – 97.9%2026-08-24source ↗
Claude Opus 4.7Anthropic86.7%76.6% – 96.7%2026-08-06source ↗
Claude Sonnet 5Anthropic80.0%68.2% – 91.8%2026-08-06source ↗
Llama 4 MaverickMeta20.6%10.9% – 30.3%2025-04-08source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.