StepFun's StepAudio 3 Realtime reasons while speaking, hits 90.6 on MMSU

stepfun-ai · hf · 2026-09-16

StepFun released the technical report for StepAudio 3 Realtime, an audio-language foundation model built around a continuous listen-converse-think-act loop.

In reasoning mode it scores 73.0 macro on StepAudioChat, 90.6 on MMSU, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and 56.0% task success on τ-Voice — matching dedicated reasoning models while speaking in real time.

Related event: StepFun Releases StepAudio 3 Family of Five Audio Models(6 posts)→

Original post →

More from Models

Models channel →