FULL STORY
Fish Audio Launches S2.1 Pro and Free API
Fish Audio launched the S2.1 Pro voice model alongside $50M in funding, later opening it as a free API optimized for single-card H200 inference across 83 languages.
2026-07-29 ~ 2026-07-30 · 2 episodes · 13 posts
Episode 1 · Fish Audio Launches S2.1 Pro Voice Model and Raises $52M (2026-07-29, 11 posts)
Fish Audio has officially released S2.1 Pro, a voice cloning model designed for production environments, alongside announcing a $52 million seed funding round. The model focuses on real-time conversation and emotional expression, enabling high-fidelity voice cloning with minimal audio samples while directly challenging competitors in both generation speed and cost.
Confirmed
- Funding & Launch: Fish Audio officially announced the completion of a $52 million seed funding round and the simultaneous public release of the S2.1 Pro voice model.
- Cloning Threshold: According to various bloggers, the model requires only 5 to 10 seconds of audio to clone a voice, preserving the original speaker's tone, rhythm, emotion, and accent.
- Performance Metrics: Official data shows S2.1 Pro has a first-audio latency of roughly 90ms, fully capable of supporting smooth real-time voice conversations. The model covers 83 languages and allows real-time emotion and style control via tags like [whispers sweetly].
- Competitor Comparison: Fish Audio claims that S2.1 Pro's generation speed is twice that of Cartesia, while its operating cost is only one-sixth of ElevenLabs.
Why It Matters
- Lowering Development Barriers: The 5-second cloning sample requirement and ultra-low 90ms latency drastically reduce the development and integration barriers for applications like real-time voice assistants and interactive game NPCs.
- Intensifying Market Competition: By directly benchmarking against Cartesia and ElevenLabs with specific multiples in speed and cost advantages, the price and performance war in the Text-to-Speech (TTS) and voice cloning markets is set to escalate further.
- Fish Audio says its new voice model clones a speaker from 5 seconds of audio — JafarNajafov · 2026-07-29
- Fish Audio says its S2.1 Pro voice model can start in 90 ms across 83 languages — testingcatalog · 2026-07-29
- Fish Audio Launches S2.1 Pro: 90ms Latency TTS Across 83 Languages — testingcatalog · 2026-07-29
- Fish Audio says its voice clone now needs just 10 seconds of speech — aryanXmahajan · 2026-07-29
- Fish Audio Raises $52M Seed, Launches S2.1 Pro Voice Cloning Model — thisdudelikesAI · 2026-07-29
- Fish Audio launches S2.1 Pro voice model and says it raised a $52M seed round — LinusEkenstam · 2026-07-29
- Fish Audio raises $52M seed for AI voice models as ARR hits $21M — davidyin44 · 2026-07-29
- Fish Audio raises $52M seed and launches S2.1 Pro voice model publicly — thisdudelikesAI · 2026-07-29
- Fish Audio marks one year with S2.1 Pro, its most advanced voice model yet — JafarNajafov · 2026-07-29
- Fish Audio Raises $52M Seed, Launches Open-Weight Voice Model S2.1 Pro — udmrzn · 2026-07-30
- Fish Audio Makes S2.1 Pro Voice Model Free, Costs 1/6 of ElevenLabs — rohanpaul_ai · 2026-07-30
Episode 2 · Fish Audio Launches Free TTS API Optimized for Single H200 GPU (2026-07-30, 2 posts)
Fish Audio has launched its S2.1 Pro voice model as a free Text-to-Speech API, supporting 83 languages. This is enabled by extreme inference optimization, achieving 0.17 RTF and 125 audio tokens per second on a single NVIDIA H200 GPU.
- Fish Audio Details Inference Stack: 0.17 RTF on a Single H200 GPU — rohanpaul_ai · 2026-07-30
- Fish Audio Launches Free TTS API: Squeezing 125 Audio Tokens/sec on a Single H200 — rohanpaul_ai · 2026-07-30