StepAudio 3 API docs detail full model lineup from TTS to music generation
StepFun_ai · x · 2026-09-16
StepFun published API access and full documentation for the StepAudio 3 family, detailing stepaudio-3-tts, asr-max, realtime-preview (full-duplex with adaptive reasoning), gen-preview (unified speech/SFX/ambience/music generation), and music-preview, plus the earlier stepaudio-2.5 lineup. TTS/ASR support Chinese, English, Japanese, Korean, French and Spanish; realtime currently supports Chinese and English.
Related event: StepFun Launches StepAudio 3 Family of Five Audio Models(4 posts)→
More from Models
- AI2's NGU sampling fixes RL for LLMs that only improves easy tasks — allenai · 2026-09-16
- With retries and pooled selection, Qwen3.8 27B hits 92.04% on DeepSWE 1.1, ~18 pts above GPT-6 Astra — S_Conradi · 2026-09-16
- Rohan Paul: With MCP Behind Every Model, Picking One Becomes Optional — kevinkern · 2026-09-16
- Redditor Claims Cursor's Grok 4.6 Gave an Oddly Self-Aware Reply — ISmellARatt · 2026-09-16
- Jev Benchmark Launches: $42 per Billion Input Tokens, Output Free Forever — cephaloform · 2026-09-16
- Same Prompt, Four Models Behind One MCP: Only One Got It Right — rohanpaul_ai · 2026-09-16