HumanEval.org

Benchmarks

Terminal-Bench 4.0

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelTerminal-Bench 4.0ObservedSource
GPT-6 AstraOpenAI58.2%55.4% – 61.0%2026-09-03source ↗
Claude Fable 5.1Anthropic57.9%54.1% – 61.6%2026-09-01source ↗
Claude Opus 5Anthropic51.8%48.4% – 55.2%2026-07-24source ↗
Claude Fable 5Anthropic44.5%40.7% – 48.4%2026-06-09source ↗
GLM-5.3Z.ai41.8%38.6% – 45.0%2026-08-14source ↗
GPT-5.6 SolOpenAI37.3%33.5% – 41.0%2026-06-26source ↗
Claude Opus 4.8Anthropic23.6%20.1% – 27.2%2026-05-28source ↗
GPT-5.6 TerraOpenAI21.5%18.3% – 24.8%2026-06-26source ↗
Grok 4.6SpaceXAI20.3%17.2% – 23.4%2026-08-12source ↗
Gemini 3.8 FlashGoogle19.1%15.7% – 22.4%2026-09-02source ↗
GPT-5.6 LunaOpenAI17.3%14.4% – 20.1%2026-06-26source ↗
Claude Sonnet 5Anthropic12.4%9.4% – 15.5%2026-06-30source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: %, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.