Bland Speech v3 tops blind TTS benchmark with Elo 1237, second only to real humans

ycombinator · x · 2026-09-25

Voice-calling company Bland released Bland Speech v3, billed as the most human text-to-speech model. On intelligence.ai's Audio Realism Bench (blind Elo listening test), v3 scores 1237 — the best machine voice, behind only hidden real-human recordings (1384) — beating MAI-Voice-2, Grok TTS, GPT Realtime 2, Gemini TTS models, and Eleven v3. Customer Superunit reports v3 beat rivals by 27% on 19,000+ employee verification calls and made it the default voice.

Original post →

More from Models

Models channel →