StepFun's StepAudio 3 Realtime reasons while speaking, hits 90.6 on MMSU
stepfun-ai · hf · 2026-09-16
StepFun released the technical report for StepAudio 3 Realtime, an audio-language foundation model built around a continuous listen-converse-think-act loop.
- Deep Perception captures acoustic cues to interpret user intent
- Seamless Duplex models synchronized audio streams to handle pauses, backchannels, and interruptions
- Think-While-Speaking runs private reasoning in parallel with spoken delivery, resolving the deliberation-vs-latency tradeoff
- An integrated Voice Agent executes tools asynchronously without disrupting dialogue
In reasoning mode it scores 73.0 macro on StepAudioChat, 90.6 on MMSU, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and 56.0% task success on τ-Voice — matching dedicated reasoning models while speaking in real time.
Related event: StepFun Releases StepAudio 3 Family of Five Audio Models(6 posts)→
More from Models
- all-MiniLM-L6-v2 still trending on Hugging Face, its creator says stop using it — tomaarsen · 2026-09-16
- HF researcher publicly questions Claude limits: monthly exhausted but weekly 84% left — NielsRogge · 2026-09-16
- Meta's promised Muse Spark open weights are over a month late and counting — RishiFurfox · 2026-09-16
- Ex-Huawei researcher: AI is great at small-step optimization, but taste and system design remain out of reach — yangyi · 2026-09-16
- First Jev Test Shows Only 80% Agreement with Verified Gemini 3.5 Flash Workflow — mayfer · 2026-09-16
- New model Jev plays Super Mario Bros in real time on fast inference — hardimanjames · 2026-09-16