HumanEval.org

Leaderboards

Ratings come from blind human votes; confidence intervals are always shown, and models with overlapping intervals share a rank. External benchmark numbers appear on the model-type boards as context, with full provenance.

Model types

Unified model boards

One multi-metric table per model type: our HumanEval rating as the anchor, next to measured performance, specs and curated third-party benchmarks.

Arena categories

Category leaderboards

One board per battle category, computed from blind human votes with 95% confidence intervals.