Terminal-Bench 4.0
last retrieved 2026-09-08 21:40 UTC
Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.
| Model | Terminal-Bench 4.0 | Observed | Source |
|---|---|---|---|
| GPT-6 AstraOpenAI | 58.2%55.4% – 61.0% | 2026-09-03 | source ↗ |
| Claude Fable 5.1Anthropic | 57.9%54.1% – 61.6% | 2026-09-01 | source ↗ |
| Claude Opus 5Anthropic | 51.8%48.4% – 55.2% | 2026-07-24 | source ↗ |
| Claude Fable 5Anthropic | 44.5%40.7% – 48.4% | 2026-06-09 | source ↗ |
| GLM-5.3Z.ai | 41.8%38.6% – 45.0% | 2026-08-14 | source ↗ |
| GPT-5.6 SolOpenAI | 37.3%33.5% – 41.0% | 2026-06-26 | source ↗ |
| Claude Opus 4.8Anthropic | 23.6%20.1% – 27.2% | 2026-05-28 | source ↗ |
| GPT-5.6 TerraOpenAI | 21.5%18.3% – 24.8% | 2026-06-26 | source ↗ |
| Grok 4.6SpaceXAI | 20.3%17.2% – 23.4% | 2026-08-12 | source ↗ |
| Gemini 3.8 FlashGoogle | 19.1%15.7% – 22.4% | 2026-09-02 | source ↗ |
| GPT-5.6 LunaOpenAI | 17.3%14.4% – 20.1% | 2026-06-26 | source ↗ |
| Claude Sonnet 5Anthropic | 12.4%9.4% – 15.5% | 2026-06-30 | source ↗ |
— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.
Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.