Cekura speech benchmark: 2,214 live calls, GPT Realtime 2.1 most reliable, Phonic v1 fastest at 1.59s
graceisford · x · 2026-09-29
Cekura launched a speech-to-speech benchmark running the same phone agent across 8 realtime voice models (OpenAI, Google, xAI, Amazon, Phonic), 82 scenarios, 3 runs each, 2,214 live calls, with a cascade agent as reference. Only scenarios passing all 3 runs counted as reliable.
- GPT Realtime 2.1 most reliable: passed 65 of 82 scenarios on all three runs
- Phonic v1 fastest: 1.59s response latency
- Workflows included clinic booking and a longer Medicare intake, requiring correct outcomes and saved data
Takeaway: production teams are moving from cascaded stacks to speech-native models — though note this thread is amplified by Phonic itself.
More from coding & agent
- One prompt gets Opus 5.5 to write an entire space film in code, frames and music included — prasenx · 2026-09-29
- Concurrent agent memory queries melted under alert storms: 140 recall calls cut to 4, P99 latency down to 185ms — Amatul2008 · 2026-09-29
- Opus 5.5 (High) hits #2 in Agent Arena at 40-56% lower cost, new Pareto frontier — arena · 2026-09-29
- JupyterLite WebMCP puts a coding agent inside your live browser notebook — OpenAIDevs · 2026-09-29
- Faraday lets agents navigate 3D scans in-browser via WebMCP — OpenAIDevs · 2026-09-29
- WebMCP winners: Observatory lets humans and agents co-edit fantasy maps on one canvas — OpenAIDevs · 2026-09-29