HumanEval.org

Benchmarks

TTS Arena V2

last retrieved 2026-09-08 21:40 UTC

Third-party scores shown as context with provenance — our own arena ratings from blind human votes stay alongside every row and remain the primary signal.

ModelTTS Arena V2ObservedSource
CastleFlow v1.0Async AI15581538 – 15782026-09-08source ↗
Inworld TTS MAXInworld15571536 – 15782026-09-08source ↗
Papla P1Papla15471520 – 15742026-09-08source ↗
Inworld TTSInworld15391518 – 15602026-09-08source ↗
Hume OctaveHume AI15231500 – 15462026-09-08source ↗

— = not yet evaluated by HumanEval (no eligible arena votes). External values are the newest observation per model; hover a value for source notes and a rating for its 95% CI. Unit: Elo, higher is better.

Not enough overlap with our arena ratings for a correlation scatter yet — fewer than two scored models are rated.