AssemblyAI launches Universal-3.6 Pro Realtime, tops two streaming speech benchmarks
AssemblyAI · x · 2026-09-30
AssemblyAI released Universal-3.6 Pro Realtime, which leads two independent benchmarks for streaming speech recognition: it sits on the Pareto frontier of @trydaily's Piecat open STT benchmark and has the lowest word error rate on real voice-agent audio in @covaldev's live leaderboard.
It targets scenarios other benchmarks skip: 98.5% accuracy on short answers ("No", "Nah", "Nuh-uh") in noisy rooms and phone lines; background voices (TV, coworkers) kept out of transcripts; 32 languages with automatic detection and mid-sentence language switching; entity-aware endpointing that holds the turn open while a caller reads out a phone number. Trained on tens of thousands of hours of real voice-agent and telephony audio.
More from Multimodal
- AI-generated explainer videos hit stunning quality, with 'How a GPU Works' as the latest proof — Dr_Singularity · 2026-09-30
- GPT-6.1 Sol near-Astra quality in 3D work while running ~30% faster — cedric_chee · 2026-09-30
- Midjourney + Kling 4.0 Flash combo wows creators in video generation tests — gen_ericai · 2026-09-30
- UniMate open-sources unified text-to-animation model that drives diverse 3D skeletons — grandorganics · 2026-09-30
- Redditor creates original Spanish-language AI animated film with ComfyUI, Flux and MiniMax H3 — Low_Masterpiece_7861 · 2026-09-30
- fal opens its Agent to everyone: image, video, audio and 3D in one interface — adamho · 2026-09-30