AssemblyAI launches Universal-3.6 Pro Realtime, tops two streaming speech benchmarks

AssemblyAI · x · 2026-09-30

AssemblyAI released Universal-3.6 Pro Realtime, which leads two independent benchmarks for streaming speech recognition: it sits on the Pareto frontier of @trydaily's Piecat open STT benchmark and has the lowest word error rate on real voice-agent audio in @covaldev's live leaderboard.

It targets scenarios other benchmarks skip: 98.5% accuracy on short answers ("No", "Nah", "Nuh-uh") in noisy rooms and phone lines; background voices (TV, coworkers) kept out of transcripts; 32 languages with automatic detection and mid-sentence language switching; entity-aware endpointing that holds the turn open while a caller reads out a phone number. Trained on tens of thousands of hours of real voice-agent and telephony audio.

Related event: AssemblyAI Launches Universal-3.6 Pro Realtime, Tops Two Streaming Speech Benchmarks(2 posts)→

Original post →

More from Multimodal

Multimodal channel →