StepAudio 3 Realtime Technical Report
Bin Lin, Bo Zhao, Boyang Zhang, Boyong Wu, Chao Yan, Chen Geng, Chen Wu, Cheng Yi, Chengli Feng, Chenglin Zhu, Chengting Feng, Chengyuan Yao, Daijiao Liu, DanNi Wan, Daxin Jiang, Dongjian Li, Dongqing Pang, Fei Tian, Feng Tian, Future Li, Gang Yu, Guanglong Yang, Haoyang Zhang, Hongyuan Wang, Jia Peng, Jiahao Song, Jialong Xue, Jiamin Fan, Jiangjie Zhen, Jianzheng Gao, Jincheng Wen, Jinghua Liang, Jinglan Gong, Jun Chen, Li Xie, Liang Zhao, Lifang Zhang, Lingli Ji, Lun Cai, Min Xu, Peilin Li, Peng Yang, Pengfei Tan, Qingjian Lin, Qinxin Du, Ruijie Xiong, Runze Li, Shenghua Hu, Shengqian Qin, Shi Qiu, Siqi Tu, Siyi Zhou, Tianjiao Deng, Wanying Lu, Weiming Niu, Wen Sun, WenWen Qu, Xiangyu Zhang, Xianwei Zhang, Xiaosu Su, Xing Chen, Xinyu Liu, Xuerui Yang, Yan Wu, Yang Li, Yang Yang, Yechang Huang, Yibo Zhu, Yifan Zhang, Yinuo Yan, Youjun Chen, Yu Fu, Yu Luo, Yu Zhou, Yujie Chen, Yumang Wang, Yunzhou Ju, Yuxiang Yang, Yuxin Li, Yuxin Zhang, Zekai Liu, Zengwei Yao, Zhaoxin Yuan, Zhenwei Mou, Zhiquan Zhang, Zhiyue Wu, Zichao Li, Zichao Zhou, Ziqi Ren, Zixuan Wang
cs.SD, eess.AS
2026-09-12
StepFun's 196B MoE voice model thinks while it talks. It scores 98.9 on full-duplex (above Qwen) and 90.6 on MMSU; realtime StepAudioChat is 70.4, near dedicated reasoners.
A realtime voice assistant has to do three jobs at once: hear what the user said, take or yield the floor at the right moment, and still reason when the request is hard. Most stacks split those jobs. ASR listens, an LLM thinks, TTS talks. Full-duplex models can overlap listening and speaking, then stall as soon as the answer needs a long chain of thought. Tools make the timing worse. A weather lookup fits in a turn; changing a flight often outlives the sentence that asked for it.
StepAudio 3 Realtime, from StepFun, folds the split pipeline into one loop: listen, converse, think, act. The design bet is that private reasoning can run in parallel with spoken delivery instead of blocking it.
The model is a mixture-of-experts with about 196 billion parameters and 11 billion active per token. The language backbone is Step 3.7 Flash. The audio frontend is the Audio Transformer from Qwen3-Omni, mapped into the LLM by an adapter. Pretraining runs in three stages (modality alignment, mixed multimodal training, cooldown) at a 32K sequence length over 1.2T tokens, with a high share of pure text so the base model does not forget how to reason. Midtraining stretches context to 128K and raises the mix of audio understanding and voice-agent data.
Four mechanisms share the same timeline.
Capability mixing is a 3:1:1:1 weighted average of four teacher checkpoints, not a router.
ASR Max is strong on public sets. LibriSpeech test-clean WER is 1.18, under Doubao 2.0 ASR at 2.94, Seed 2.0 Lite at 1.47, and HY3.0 ASR Preview at 1.38. AISHELL-1 CER is 0.49, the best of the four. WenetSpeech test-net 3.99 and test-meeting 4.35 trail HY3.0's 3.71 and 4.12 by a little. On ContextASR-Bench with no hotword injection, English macro error is 5.67% and Mandarin 1.23%, versus 6.60% and 1.69% for HY3.0. Those numbers describe the ASR specialist, not the realtime talker.
Audio understanding averages 81.3 across eight benches, 0.5 behind Gemini 3.1 Pro at 81.8. MMSU 90.6 vs 83.6 and MMAR 86.5 vs 81.7 are the widest leads. AudioMultiChallenge is 49.3 against Gemini 3.1 Pro's 67.0; multi-turn constraint tracking is the hole. Cutting SFT from about 2 million random examples to about 100K high-quality ones lifts MMSU from 78.78 to 89.70 and MMAR from 74.70 to 84.50.
On the Artificial Analysis full-duplex subset the Overall score is 98.9, above Qwen Audio 3.0 Realtime Plus at 98.4 and GPT-realtime-2 High at 95.3. Turn taking is 100.0, interruptions 99.0, pauses 98.9, backchannels 98.0.
| Setting | Metric | Result |
| ASR Max vs HY3.0 | LibriSpeech test-clean WER | 1.18 vs 1.38 |
| Realtime vs Gemini 3.1 Pro | MMSU | 90.6 vs 83.6 |
| Realtime vs Qwen Audio 3.0 RT Plus | Full-Duplex Overall | 98.9 vs 98.4 |
| Realtime vs Grok Voice | τ-Voice macro success | 56.0% vs 56.5% |
StepAudioChat is a closed text benchmark that isolates dialogue skill from prosody and barge-in. Reasoning mode scores 73.0 macro, above Doubao 70.5 and DeepSeek-V4-Flash 71.4, below Kimi K3 at 77.1. The realtime Think-While-Speaking system scores 70.4, almost even with Doubao's reasoning mode, while instruction following falls from 66.3 to 54.1. On τ-Voice, telecom is 70.2, the best of the four; retail is 37.7 against Grok's 49.7. HMMT February 2026 is 86.8 versus Gemini 3 Flash at 85.9; GPQA Diamond is 83.0 versus 90.3.
Adaptive Thinking cuts explicit-think rates to 51.5%–82.0%, and the Reasoning score drops from 71.89 with full thinking to 66.80. Fewer think turns do not mean the budget landed on the turns that need it. MTP3 with typical acceptance is about 2.05× wall-clock; the fifth head's strict marginal accept rate is 5.7%, so extra heads taper off fast.
The product question is whether a voice model can think while it talks without treating every "yeah" as an interrupt. This report turns that into a dual-call schedule and still lands dialogue scores close to dedicated reasoning models. 98.9 full-duplex and 90.6 MMSU are numbers a competitor can actually score against.
It is a technical report, not a weight release. Teacher merging, a closed dialogue bench, and a misallocated Adaptive Thinking policy all say the same thing: the system is competitive and still being tuned. Instruction following losing more than 12 points in realtime mode is the number a customer-support team should watch before the duplex headline.
The paper flags the 17.7-point AudioMultiChallenge gap, weak retail tool use, and uneven Adaptive Thinking. Realtime instruction following drops more than 12 points; speaking while thinking costs obedience.
StepAudioChat cannot be audited from outside. Full-duplex and τ-Voice follow Artificial Analysis's implementation. ASR numbers belong to ASR Max. Serving cost, time-to-first-audio, and concurrency for a 196B MoE are almost absent. The merged dialogue macro of 73.0 still sits under the best teacher's 74.2; averaging parameters buys balance, not a win on every axis.