Fish Audio says its S2.1 Pro voice model can start in 90 ms across 83 languages
testingcatalog · x · 2026-07-29
Fish Audio is promoting S2.1 Pro as a real-time voice model with low latency, broad language coverage, and inline emotion control.
- The model reportedly reaches about 90 ms time-to-first-audio.
- It supports 83 languages and free-form bracket tags such as [whispers sweetly] or [laughing nervously] for emotion control.
- A free tier uses the same model at no cost for testing.
- The product page positions it for voiceovers, audiobooks, customer support, and other expressive audio workflows.
Related event: Fish Audio Launches S2.1 Pro Voice Model with Ultra-Low Latency(4 posts)→
More from Multimodal
- Claude 5 Opus generates a textureless dirt-road car demo entirely on its own — ChrisGPT · 2026-07-29
- TILT improves compositional text-to-image generation with a model-intrinsic reward — Debottam Dutta · 2026-07-29
- Claude 5 Opus turns a no-texture dirt-road car demo into fully generated game graphics — ChrisGPT · 2026-07-29
- Stream3D turns frozen 3D generators into streaming models with bounded memory — pliang279 · 2026-07-29
- Creator turns Agent One into a 90-second cinematic horror trailer — LudovicCreator · 2026-07-29
- AI image experiment moved from realism to silkscreen after moiré issues — emollick · 2026-07-29