StepFun launches StepAudio 3: five audio models topping realtime voice leaderboards

StepFun_ai · x · 2026-09-16

StepFun released StepAudio 3, a family of five audio models covering realtime voice, ASR, speech generation, audio scenes, and music. The realtime model ranks #1 on Artificial Analysis for Conversational Dynamics (98.9%) and Speech Reasoning (99.7%), and ASR hits 1.7% WER, matching the leaderboard best. Models support interruptions, reasoning while speaking, and tool calls for voice agents, with multilingual support and API availability now.

Related event: StepFun Launches StepAudio 3 Family of Five Audio Models(4 posts)→

Original post →

More from Models

Models channel →