StepAudio 3 API docs detail full model lineup from TTS to music generation

StepFun_ai · x · 2026-09-16

StepFun published API access and full documentation for the StepAudio 3 family, detailing stepaudio-3-tts, asr-max, realtime-preview (full-duplex with adaptive reasoning), gen-preview (unified speech/SFX/ambience/music generation), and music-preview, plus the earlier stepaudio-2.5 lineup. TTS/ASR support Chinese, English, Japanese, Korean, French and Spanish; realtime currently supports Chinese and English.

Related event: StepFun Launches StepAudio 3 Family of Five Audio Models(4 posts)→

Original post →

More from Models

Models channel →