GPT-5.6 Solnot yet evaluated by HumanEval
HumanEval ratings
Not yet evaluated by HumanEval — this model is listed for context from external benchmarks and has not collected blind human votes in the arena. Our rating appears here once it joins the roster and gets eligible votes.
External benchmarks
Curated third-party results, shown as context with provenance — the HumanEval rating from blind human votes remains the verdict.
General Intelligence
Reasoning
| LiveBench — Reasoning | 91.7 | 2026-09-07 | LiveBench |
| GPQA Diamond | 93.5%90.4% – 96.6% | 2026-07-09 | Epoch AI Benchmarking Hub |
Coding
| LiveBench — Coding | 83.9 | 2026-09-07 | LiveBench |
Mathematics
| LiveBench — Mathematics | 96.2 | 2026-09-07 | LiveBench |
| OTIS Mock AIME 2024–2025 | 100.0%100.0% – 100.0% | 2026-07-09 | Epoch AI Benchmarking Hub |
Agentic
| LiveBench — Agentic Coding | 56.2 | 2026-09-07 | LiveBench |
| Terminal-Bench 4.0 | 37.3%33.5% – 41.0% | 2026-06-26 | Terminal-Bench |
Computer use
| OSWorld 2.0 (binary accuracy) | 27.3% | 2026-09-03 | OSWorld |
Language
| LiveBench — Language | 87.7 | 2026-09-07 | LiveBench |
Data analytics
| LiveBench — Data Analysis | 79.8 | 2026-09-07 | LiveBench |
Instruction following
| LiveBench — Instruction Following | 71.8 | 2026-09-07 | LiveBench |
Benchmark percentile by metric group
Percentile of this model within the 19 listed LLM models per metric group (100 = best, averaged across the group's external benchmarks), against the roster median. External data only — not the HumanEval rating.