Third-party benchmark: AssemblyAI Universal 3.5 Pro tops 15 STT models at 1.93% WER and 489ms latency
AssemblyAI · x · 2026-09-24
The third-party Converse-STT benchmark (Cekura Bench) ran 15 speech-to-text models through the same audio pipeline, and AssemblyAI Universal 3.5 Pro came first on both accuracy and latency.
- On 1,000 read-aloud clips from Pipecat FLEURS, it posted the lowest word error rate at 1.93%, ahead of Google Chirp 3 (2.00%), Reson8 (2.10%) and GPT Realtime Whisper (2.16%).
- Its median time to first text (TTFT P50) was 489ms, the fastest of any model tested — Google Chirp 3 took 4.48s on the same metric.
- Overall it sits on the quality-speed Pareto frontier, making it a strong pick for latency-sensitive voice agent pipelines.
Full results are public on Cekura Bench, with source code available.
More from Models
- Users find Opus 5.5 has remarkably strong hearing, works via spectrograms — repligate · 2026-09-24
- GPT-6 Astra agent plays Left 4 Dead 2 in first real-time run — imjustnewatai · 2026-09-24
- Same Prompt, Opposite Results: GPT-4 Goes Silent 30/30 Where GPT-3.5 Never Stops — rayanpal_ · 2026-09-24
- Box says Opus 5.5 cuts token usage 63% and runs 30% faster than Opus 5 — bcherny · 2026-09-24
- Claude Opus 5.5 Max Tops Code Arena WebDev at 1818, Leading GPT-6 Astra by 26 — airesearch12 · 2026-09-24
- 36B Agent Model Runs Fully Local on Snapdragon X2 Elite With Just 32GB RAM — Kyrannio · 2026-09-24