OTIS Mock AIME 2024–2025
last retrieved 2026-09-08 21:40 UTC
Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.
| Model | OTIS Mock AIME 2024–2025 | Observed | Source |
|---|---|---|---|
| Claude Fable 5.1Anthropic | 100.0%100.0% – 100.0% | 2026-09-01 | source ↗ |
| GPT-5.6 SolOpenAI | 100.0%100.0% – 100.0% | 2026-07-09 | source ↗ |
| GPT-6 AstraOpenAI | 100.0%100.0% – 100.0% | 2026-08-30 | source ↗ |
| Claude Fable 5Anthropic | 99.7%99.2% – 100.3% | 2026-06-10 | source ↗ |
| GPT-5.6 TerraOpenAI | 99.7%99.2% – 100.3% | 2026-07-09 | source ↗ |
| Qwen3.8 Max (0902)Alibaba Qwen | 99.4%98.7% – 100.2% | 2026-08-04 | source ↗ |
| Grok 4.6SpaceXAI | 99.2%98.3% – 100.1% | 2026-08-14 | source ↗ |
| Claude Opus 5Anthropic | 98.9%96.7% – 101.1% | 2026-07-24 | source ↗ |
| DeepSeek V4 Pro (0813)DeepSeek | 98.6%96.7% – 100.5% | 2026-08-18 | source ↗ |
| Claude Opus 4.8Anthropic | 98.3%95.6% – 101.1% | 2026-06-07 | source ↗ |
| GPT-5.6 LunaOpenAI | 98.3%96.0% – 100.6% | 2026-07-09 | source ↗ |
| Kimi K3Moonshot AI | 97.2%95.0% – 99.4% | 2026-07-16 | source ↗ |
| Gemini 3.1 Pro PreviewGoogle | 95.6%89.5% – 101.6% | 2026-08-06 | source ↗ |
| GLM-5.3Z.ai | 91.1%84.3% – 97.9% | 2026-08-24 | source ↗ |
| Claude Opus 4.7Anthropic | 86.7%76.6% – 96.7% | 2026-08-06 | source ↗ |
| Claude Sonnet 5Anthropic | 80.0%68.2% – 91.8% | 2026-08-06 | source ↗ |
| Llama 4 MaverickMeta | 20.6%10.9% – 30.3% | 2025-04-08 | source ↗ |
— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.
Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.