Benchmarks
Curated third-party benchmark results, shown as context with provenance — every score links to its source with the date it was observed and the date we retrieved it. Our own blind human-vote arena ratings remain the primary signal on this site; external numbers never feed them. See methodology for why we treat published benchmarks with caution.
Sources
last retrieved 2026-09-08 21:40 UTC
Terms of Use (lmarena.ai/terms-of-use, read 2026-09-08) §5 forbid reproducing/mirroring/commercially exploiting the Service; they do not address quoting individual scores. We quote single Arena Scores with attribution + link as factual context (no mirror, no bulk redistribution). FLAGGED FOR OPERATOR REVIEW — set active=false to hide every LMArena number.
last retrieved 2026-09-08 21:40 UTC
Community arena (github.com/TTS-AGI/TTS-Arena, Apache-2.0) exposing a public /api/leaderboard JSON; no explicit data licence — quoted with attribution. Stealth/anonymous entries are excluded. ± = the API's "uncertainty" field. FLAGGED FOR OPERATOR REVIEW.
last retrieved 2026-09-08 21:40 UTC
Site footer: "This website is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License" (CC BY-SA 4.0). Category values are computed from the published per-task table (table_2026_06_25.csv) with the site's stated formula (category = mean of its task scores; Overall = mean of category averages).
last retrieved 2026-09-08 21:40 UTC
benchmark_data.zip README (2026-09): "Epoch AI's data is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons Attribution license" (CC BY 4.0). Cite: Epoch AI, "AI Benchmarking Hub", https://epoch.ai/benchmarks. Values are Epoch's own runs ("Best score (across scorers)"); 95% CI = ±1.96·stderr.
last retrieved 2026-09-08 21:40 UTC
Public leaderboard page without an explicit data licence; individual accuracies quoted with attribution. Cross-checked identical against Epoch AI's CC BY 4.0 redistribution (hle_external.csv). Displayed ± is the standard error; CI = ±1.96·SE.
last retrieved 2026-09-08 21:40 UTC
Leaderboard data + evaluation code are MIT-licensed (github.com/swe-bench/experiments); values quoted with attribution. Only entries marked "checked" by the SWE-bench team are imported.
last retrieved 2026-09-08 21:40 UTC
Official Terminal-Bench 4.0 leaderboard (tbench.ai); Epoch AI records the leaderboard data as Apache-2.0. Values depend on the agent harness used (recorded in notes); ± = 95% CI half-width as published.
last retrieved 2026-09-08 21:40 UTC
Official results JSON published by the OSWorld team (project code Apache-2.0); also redistributed by Epoch AI under CC BY 4.0. Binary accuracy on the full OSWorld 2.0 task set, 500-step budget, batch-tool setting.
Catalogue
| Benchmark | Source | Metric group | Unit | Better | Models | Last observed |
|---|---|---|---|---|---|---|
| LMArena Text Arena (overall) | LMArena | General Intelligence | Elo | higher | 18 | 2026-09-02 |
| LMArena Text-to-Image Arena | LMArena | Image quality (human preference) | Elo | higher | 5 | 2026-09-07 |
| LMArena Text-to-Video Arena | LMArena | Video quality (human preference) | Elo | higher | 5 | 2026-09-04 |
| TTS Arena V2 | TTS Arena (V2) | Speech quality (human preference) | Elo | higher | 5 | 2026-09-08 |
| LiveBench — Overall | LiveBench | General Intelligence | score | higher | 18 | 2026-09-07 |
| LiveBench — Reasoning | LiveBench | Reasoning | score | higher | 18 | 2026-09-07 |
| GPQA Diamond | Epoch AI Benchmarking Hub | Reasoning | % | higher | 16 | 2026-08-30 |
| Humanity's Last Exam | Scale AI SEAL — Humanity's Last Exam | Reasoning | % | higher | 5 | 2026-09-08 |
| LiveBench — Coding | LiveBench | Coding | score | higher | 18 | 2026-09-07 |
| SWE-bench Verified | SWE-bench | Coding | % | higher | 1 | 2025-07-20 |
| LiveBench — Mathematics | LiveBench | Mathematics | score | higher | 18 | 2026-09-07 |
| OTIS Mock AIME 2024–2025 | Epoch AI Benchmarking Hub | Mathematics | % | higher | 17 | 2026-09-01 |
| LiveBench — Agentic Coding | LiveBench | Agentic | score | higher | 18 | 2026-09-07 |
| Terminal-Bench 4.0 | Terminal-Bench | Agentic | % | higher | 12 | 2026-09-03 |
| OSWorld 2.0 (binary accuracy) | OSWorld | Computer use | % | higher | 4 | 2026-09-03 |
| LiveBench — Language | LiveBench | Language | score | higher | 18 | 2026-09-07 |
| LiveBench — Data Analysis | LiveBench | Data analytics | score | higher | 18 | 2026-09-07 |
| LiveBench — Instruction Following | LiveBench | Instruction following | score | higher | 18 | 2026-09-07 |
Scores enter this catalogue through curated imports and manual review only — there are no automated scrapers. A benchmark's numbers are the newest observation per model as reported by its source; historical observations are kept and shown on each benchmark's page.