Cekura speech benchmark: 2,214 live calls, GPT Realtime 2.1 most reliable, Phonic v1 fastest at 1.59s

graceisford · x · 2026-09-29

Cekura launched a speech-to-speech benchmark running the same phone agent across 8 realtime voice models (OpenAI, Google, xAI, Amazon, Phonic), 82 scenarios, 3 runs each, 2,214 live calls, with a cascade agent as reference. Only scenarios passing all 3 runs counted as reliable.

Takeaway: production teams are moving from cascaded stacks to speech-native models — though note this thread is amplified by Phonic itself.

Original post →

More from coding & agent

coding & agent channel →