Fish Audio open-sources S2, a controllable TTS model trained on 10M hours of audio

alex_verem · x · 2026-07-29

Fish Audio has open-sourced S2, a text-to-speech model focused on fine-grained control of prosody and emotion.

S2 supports natural-language inline directives such as [laugh], [whispers], and open-ended descriptions like [professional broadcast tone] placed at specific words or phrases. The system was trained on more than 10 million hours of audio across roughly 50 languages and combines reinforcement-learning alignment with a dual-autoregressive architecture. The release includes model weights, fine-tuning code, and an SGLang-based streaming inference engine, and claims state-of-the-art results against Seed-TTS, MiniMax-Speech, and even gpt-4o-mini-tts in the reported benchmarks.

Original post →

More from Multimodal

Multimodal channel →