HumanEval.org

Benchmarks

Curated third-party benchmark results, shown as context with provenance — every score links to its source with the date it was observed and the date we retrieved it. Our own blind human-vote arena ratings remain the primary signal on this site; external numbers never feed them. See methodology for why we treat published benchmarks with caution.

Sources
8
Benchmarks
18
Models covered
34
of 34 listed models
Grid coverage
38%
232 of 612 model × benchmark cells

Sources

LMArena3 benchmarks

last retrieved 2026-09-08 21:40 UTC

Terms of Use (lmarena.ai/terms-of-use, read 2026-09-08) §5 forbid reproducing/mirroring/commercially exploiting the Service; they do not address quoting individual scores. We quote single Arena Scores with attribution + link as factual context (no mirror, no bulk redistribution). FLAGGED FOR OPERATOR REVIEW — set active=false to hide every LMArena number.

last retrieved 2026-09-08 21:40 UTC

Community arena (github.com/TTS-AGI/TTS-Arena, Apache-2.0) exposing a public /api/leaderboard JSON; no explicit data licence — quoted with attribution. Stealth/anonymous entries are excluded. ± = the API's "uncertainty" field. FLAGGED FOR OPERATOR REVIEW.

LiveBench8 benchmarks

last retrieved 2026-09-08 21:40 UTC

Site footer: "This website is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License" (CC BY-SA 4.0). Category values are computed from the published per-task table (table_2026_06_25.csv) with the site's stated formula (category = mean of its task scores; Overall = mean of category averages).

last retrieved 2026-09-08 21:40 UTC

benchmark_data.zip README (2026-09): "Epoch AI's data is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons Attribution license" (CC BY 4.0). Cite: Epoch AI, "AI Benchmarking Hub", https://epoch.ai/benchmarks. Values are Epoch's own runs ("Best score (across scorers)"); 95% CI = ±1.96·stderr.

last retrieved 2026-09-08 21:40 UTC

Public leaderboard page without an explicit data licence; individual accuracies quoted with attribution. Cross-checked identical against Epoch AI's CC BY 4.0 redistribution (hle_external.csv). Displayed ± is the standard error; CI = ±1.96·SE.

SWE-bench1 benchmark

last retrieved 2026-09-08 21:40 UTC

Leaderboard data + evaluation code are MIT-licensed (github.com/swe-bench/experiments); values quoted with attribution. Only entries marked "checked" by the SWE-bench team are imported.

last retrieved 2026-09-08 21:40 UTC

Official Terminal-Bench 4.0 leaderboard (tbench.ai); Epoch AI records the leaderboard data as Apache-2.0. Values depend on the agent harness used (recorded in notes); ± = 95% CI half-width as published.

OSWorld1 benchmark

last retrieved 2026-09-08 21:40 UTC

Official results JSON published by the OSWorld team (project code Apache-2.0); also redistributed by Epoch AI under CC BY 4.0. Binary accuracy on the full OSWorld 2.0 task set, 500-step budget, batch-tool setting.

Catalogue

BenchmarkSourceMetric groupUnitBetterModelsLast observed
LMArena Text Arena (overall)LMArenaGeneral IntelligenceElohigher182026-09-02
LMArena Text-to-Image ArenaLMArenaImage quality (human preference)Elohigher52026-09-07
LMArena Text-to-Video ArenaLMArenaVideo quality (human preference)Elohigher52026-09-04
TTS Arena V2TTS Arena (V2)Speech quality (human preference)Elohigher52026-09-08
LiveBench — OverallLiveBenchGeneral Intelligencescorehigher182026-09-07
LiveBench — ReasoningLiveBenchReasoningscorehigher182026-09-07
GPQA DiamondEpoch AI Benchmarking HubReasoning%higher162026-08-30
Humanity's Last ExamScale AI SEAL — Humanity's Last ExamReasoning%higher52026-09-08
LiveBench — CodingLiveBenchCodingscorehigher182026-09-07
SWE-bench VerifiedSWE-benchCoding%higher12025-07-20
LiveBench — MathematicsLiveBenchMathematicsscorehigher182026-09-07
OTIS Mock AIME 2024–2025Epoch AI Benchmarking HubMathematics%higher172026-09-01
LiveBench — Agentic CodingLiveBenchAgenticscorehigher182026-09-07
Terminal-Bench 4.0Terminal-BenchAgentic%higher122026-09-03
OSWorld 2.0 (binary accuracy)OSWorldComputer use%higher42026-09-03
LiveBench — LanguageLiveBenchLanguagescorehigher182026-09-07
LiveBench — Data AnalysisLiveBenchData analyticsscorehigher182026-09-07
LiveBench — Instruction FollowingLiveBenchInstruction followingscorehigher182026-09-07

Scores enter this catalogue through curated imports and manual review only — there are no automated scrapers. A benchmark's numbers are the newest observation per model as reported by its source; historical observations are kept and shown on each benchmark's page.