StepFun drops five StepAudio3 models, claiming three Artificial Analysis world No.1s

新智元 · wechat · 2026-09-15

StepFun launched the StepAudio3 family on Sept 15 — five models covering TTS, full audio generation (Gen), realtime voice interaction, ASR, and music — the first full-stack audio lineup from a Chinese AI company.

It claims three Artificial Analysis world No.1s: Realtime scored 98.9% on Conversational Dynamics and 99.7% on Speech Reasoning; ASR hit a 1.7% WER, near-professional-stenographer territory. ASR combines acoustic modeling with LLM contextual reasoning and handles dialects, code-switching, whispers, and singing.

Highlights from hands-on tests: TTS naturally reproduces hesitations, self-corrections, and laughter; Realtime survived four rapid mid-sentence interruptions, running tool calls in parallel with speech generation; Gen produces complete sound scenes (dialogue, foley, ambience, score) from one prompt with per-character emotional arcs; Music accepts stems, hummed melodies, and prompts down to key, BPM, instrumentation, and vocal aesthetics, and can arrange a raw a cappella into a finished track. The piece argues audio AI competition is shifting from single-skill contests to full-stack system play.

Original post →

More from Multimodal

Multimodal channel →