Fish Audio open-sources S2, a controllable TTS model trained on 10M hours of audio
alex_verem · x · 2026-07-29
Fish Audio has open-sourced S2, a text-to-speech model focused on fine-grained control of prosody and emotion.
S2 supports natural-language inline directives such as [laugh], [whispers], and open-ended descriptions like [professional broadcast tone] placed at specific words or phrases. The system was trained on more than 10 million hours of audio across roughly 50 languages and combines reinforcement-learning alignment with a dual-autoregressive architecture. The release includes model weights, fine-tuning code, and an SGLang-based streaming inference engine, and claims state-of-the-art results against Seed-TTS, MiniMax-Speech, and even gpt-4o-mini-tts in the reported benchmarks.
More from Multimodal
- Claude 5 Opus generates a textureless dirt-road car demo entirely on its own — ChrisGPT · 2026-07-29
- TILT improves compositional text-to-image generation with a model-intrinsic reward — Debottam Dutta · 2026-07-29
- Claude 5 Opus turns a no-texture dirt-road car demo into fully generated game graphics — ChrisGPT · 2026-07-29
- Stream3D turns frozen 3D generators into streaming models with bounded memory — pliang279 · 2026-07-29
- Creator turns Agent One into a 90-second cinematic horror trailer — LudovicCreator · 2026-07-29
- AI image experiment moved from realism to silkscreen after moiré issues — emollick · 2026-07-29