StepFun launches StepAudio 3: five audio models, #1 in realtime speech evals
StepFun_ai · x · 2026-09-16
StepFun has released StepAudio 3, a family of five audio models covering realtime voice, speech recognition, speech generation, audio scene generation, and music.
- Designed for voice agents that handle interruptions, reason while speaking, and call tools
- Ranks #1 on Artificial Analysis for Conversational Dynamics (98.9%) and Speech Reasoning (99.7%)
- ASR hits 1.7% WER, matching the best result on the leaderboard
- Available now on the StepFun platform
Related event: StepFun Launches StepAudio 3 Family of Five Audio Models(4 posts)→
More from Multimodal
- Viral Storm Dance AI trend built on Midjourney + Seedance 2.5, workflow prompt shared — miilesus · 2026-09-16
- Storm dance trend recipe: Midjourney visuals + Seedance 2.5 video, prompt shared — miilesus · 2026-09-16
- Midjourney 'Pegasus' spotted in early user post with generated samples — chrisfirst · 2026-09-16
- StepFun launches StepAudio 3: five audio models topping realtime voice leaderboards — StepFun_ai · 2026-09-16
- Poolday raises $11M for a video agent that edits, assembles and QAs itself — azed_ai · 2026-09-16
- AI crow videos fool 95% of Facebook commenters as detection gets impossible — PhenomenalKid · 2026-09-16