LiveBench — Agentic Coding
last retrieved 2026-09-08 21:40 UTC
Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.
| Model | LiveBench — Agentic Coding | Observed | Source |
|---|---|---|---|
| Claude Fable 5.1Anthropic | 66.1 | 2026-09-07 | source ↗ |
| Claude Opus 5Anthropic | 65.2 | 2026-09-07 | source ↗ |
| Qwen3.8 Max (0902)Alibaba Qwen | 64.7 | 2026-09-07 | source ↗ |
| Claude Fable 5Anthropic | 62.2 | 2026-09-07 | source ↗ |
| Kimi K3Moonshot AI | 62.2 | 2026-09-07 | source ↗ |
| GLM-5.3Z.ai | 60.9 | 2026-09-07 | source ↗ |
| Claude Sonnet 5Anthropic | 59.4 | 2026-09-07 | source ↗ |
| GPT-6 AstraOpenAI | 57.3 | 2026-09-07 | source ↗ |
| Grok 4.6SpaceXAI | 57.0 | 2026-09-07 | source ↗ |
| GPT-5.6 SolOpenAI | 56.2 | 2026-09-07 | source ↗ |
| DeepSeek V4 Pro (0813)DeepSeek | 55.0 | 2026-09-07 | source ↗ |
| GPT-5.6 TerraOpenAI | 55.0 | 2026-09-07 | source ↗ |
| Gemini 3.8 FlashGoogle | 54.2 | 2026-09-07 | source ↗ |
| GPT-5.4OpenAI | 53.8 | 2026-09-07 | source ↗ |
| Claude Opus 4.7Anthropic | 50.7 | 2026-09-07 | source ↗ |
| Claude Opus 4.8Anthropic | 50.5 | 2026-09-07 | source ↗ |
| GPT-5.6 LunaOpenAI | 48.4 | 2026-09-07 | source ↗ |
| Gemini 3.1 Pro PreviewGoogle | 44.1 | 2026-09-07 | source ↗ |
— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: score, higher is better.
Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.