StepFun drops five StepAudio3 models, claiming three Artificial Analysis world No.1s
新智元 · wechat · 2026-09-15
StepFun launched the StepAudio3 family on Sept 15 — five models covering TTS, full audio generation (Gen), realtime voice interaction, ASR, and music — the first full-stack audio lineup from a Chinese AI company.
It claims three Artificial Analysis world No.1s: Realtime scored 98.9% on Conversational Dynamics and 99.7% on Speech Reasoning; ASR hit a 1.7% WER, near-professional-stenographer territory. ASR combines acoustic modeling with LLM contextual reasoning and handles dialects, code-switching, whispers, and singing.
Highlights from hands-on tests: TTS naturally reproduces hesitations, self-corrections, and laughter; Realtime survived four rapid mid-sentence interruptions, running tool calls in parallel with speech generation; Gen produces complete sound scenes (dialogue, foley, ambience, score) from one prompt with per-character emotional arcs; Music accepts stems, hummed melodies, and prompts down to key, BPM, instrumentation, and vocal aesthetics, and can arrange a raw a cappella into a finished track. The piece argues audio AI competition is shifting from single-skill contests to full-stack system play.
More from Multimodal
- Agent-searchable API serves 1.49M public-domain historical illustrations with sources — vanstriendaniel · 2026-09-15
- Synthesia rolls out Avatar Builder to all users for custom AI avatars — synthesiaIO · 2026-09-15
- Runway MCP prompt generates 30s photoreal stealth-game footage, fully shared — tlakomy · 2026-09-15
- AI music videos should be judged by how easy they are to fix, not first-draft quality — Positive-Page1631 · 2026-09-15
- MiniMax H3 T2V generates similar faces across seeds while WAN stays more varied — ThrowAwayBiCall911 · 2026-09-15
- Fixing MiniMax H3 face distortion with a portrait-distance prompt constraint — Devajyoti1231 · 2026-09-15