HumanEval.org

Benchmarks

LMArena Text-to-Image Arena

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelLMArena Text-to-Image ArenaObservedSource
GPT Image 2OpenAI13811377 – 13852026-09-07source ↗
MAI-Image-2.6Microsoft AI13311324 – 13382026-09-07source ↗
Reve 2.1Reve13011293 – 13092026-09-07source ↗
Muse ImageMeta12771271 – 12832026-09-07source ↗
Gemini 3.1 Flash Image (Nano Banana 2)Google12611256 – 12662026-09-07source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: Elo, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.