StepFun launches StepAudio 3: five audio models topping realtime voice leaderboards
StepFun_ai · x · 2026-09-16
StepFun released StepAudio 3, a family of five audio models covering realtime voice, ASR, speech generation, audio scenes, and music. The realtime model ranks #1 on Artificial Analysis for Conversational Dynamics (98.9%) and Speech Reasoning (99.7%), and ASR hits 1.7% WER, matching the leaderboard best. Models support interruptions, reasoning while speaking, and tool calls for voice agents, with multilingual support and API availability now.
Related event: StepFun Launches StepAudio 3 Family of Five Audio Models(4 posts)→
More from Models
- Predictions for a Huge AI Week: Opus 5.2, Codex Bot, and a Cheaper Sol — daniel_mac8 · 2026-09-16
- Cartesia's new voice model family draws praise; WER alone can't capture context-correct speech — buckymoore · 2026-09-16
- Replication of no-CoT evals shows GPT-Astra makes a qualitative jump across all datasets — dhadfieldmenell · 2026-09-16
- Anthropic's Astra tops spend while OpenAI's Luna dominates token usage by a lot — gdb · 2026-09-16
- DeepSeekMath-V2 makes verification the product, scaling verifier compute ahead of the generator — le_james94 · 2026-09-16
- ChatGPT Plus Work Projects Bug Persists for Days While OpenAI Marks It Resolved — OnwardUpwardForward · 2026-09-16