Leaderboards
Ratings come from blind human votes; confidence intervals are always shown, and models with overlapping intervals share a rank. External benchmark numbers appear on the model-type boards as context, with full provenance.
Model types
Unified model boards
One multi-metric table per model type: our HumanEval rating as the anchor, next to measured performance, specs and curated third-party benchmarks.
Arena categories
Category leaderboards
One board per battle category, computed from blind human votes with 95% confidence intervals.