HumanEval.org

Humans judging AI, in the open

Blind head-to-head battles, votes from real people, and ratings with confidence intervals you can reproduce from published data.

81 votes cast across 143 battles · latest ratings 2026-09-04 18:45 UTC

Top rated by category

All leaderboards →
Human-Like Chat2026-09-04 18:45 UTC
  1. 1Mock Alpha1138.8 1057.7 – 1235.4
  2. 1Mock Premium1001.2 936.9 – 1077.2
  3. 3Mock Beta860.0 785.8 – 927.7
Stealth Text2026-09-04 18:45 UTC
  1. Mock Alphaprovisional1874.0 1673.5 – 2855.4
  2. Mock Premiumprovisional1536.0 1000.0 – 1671.1
  3. Mock Betaprovisional-410.1 -877.2 – -301.9

Biggest movers

Newest models

Changelog

Full changelog →

Stay in the loop

Launch notes, new models and methodology changes — straight to your inbox, nothing else.