Fish Audio Launches S2.1 Pro: 90ms Latency TTS Across 83 Languages
testingcatalog · x · 2026-07-29
Fish Audio has launched S2.1 Pro, a production-ready text-to-speech model optimized for real-time conversational voice.
Key features:
- Ultra-low latency: 90ms time-to-first-audio, enabling natural turn-taking in live dialogue.
- Multilingual: A single voice identity covers 83 languages.
- Inline emotion control: Supports free-form bracket tags (e.g., [whispers sweetly]) for dynamic emotion steering.
- Voice cloning: Captures tone and style from just 10-30 seconds of reference audio.
- Developer-friendly: Offers a free tier and integrates into agent workflows via MCP and agent-skill support.
Related event: Fish Audio Launches S2.1 Pro Voice Model with Ultra-Low Latency(4 posts)→
More from Models
- Leak says GPT-6 slips to early September as Anthropic tests Fable 5.1 — soumitrashukla9 · 2026-07-29
- GPT-5.6 Sol Ultra finds a critical bug, then refuses to show it — haltakov · 2026-07-29
- Kimi K3 tops a benchmark chart in a repost claiming it beats Anthropic models — JarnoDuursma · 2026-07-29
- User Reports Grok's Generation Capabilities Have Gotten 'Real Cracked' — djcows · 2026-07-29
- User asks Anthropic not to deprecate Opus 4.6 until the model is fixed — oyacaro · 2026-07-29
- Hidden Trick: Manually Invoke Older Opus Models in Claude Code — voooooogel · 2026-07-29