FULL STORY

Fish Audio Launches S2.1 Pro and Free API

Fish Audio launched the S2.1 Pro voice model alongside $50M in funding, later opening it as a free API optimized for single-card H200 inference across 83 languages.

2026-07-29 ~ 2026-07-30 · 2 episodes · 13 posts

Episode 1 · Fish Audio Launches S2.1 Pro Voice Model and Raises $52M (2026-07-29, 11 posts)

Fish Audio has officially released S2.1 Pro, a voice cloning model designed for production environments, alongside announcing a $52 million seed funding round. The model focuses on real-time conversation and emotional expression, enabling high-fidelity voice cloning with minimal audio samples while directly challenging competitors in both generation speed and cost.

Confirmed

  • Funding & Launch: Fish Audio officially announced the completion of a $52 million seed funding round and the simultaneous public release of the S2.1 Pro voice model.
  • Cloning Threshold: According to various bloggers, the model requires only 5 to 10 seconds of audio to clone a voice, preserving the original speaker's tone, rhythm, emotion, and accent.
  • Performance Metrics: Official data shows S2.1 Pro has a first-audio latency of roughly 90ms, fully capable of supporting smooth real-time voice conversations. The model covers 83 languages and allows real-time emotion and style control via tags like [whispers sweetly].
  • Competitor Comparison: Fish Audio claims that S2.1 Pro's generation speed is twice that of Cartesia, while its operating cost is only one-sixth of ElevenLabs.

Why It Matters

  • Lowering Development Barriers: The 5-second cloning sample requirement and ultra-low 90ms latency drastically reduce the development and integration barriers for applications like real-time voice assistants and interactive game NPCs.
  • Intensifying Market Competition: By directly benchmarking against Cartesia and ElevenLabs with specific multiples in speed and cost advantages, the price and performance war in the Text-to-Speech (TTS) and voice cloning markets is set to escalate further.

Episode 2 · Fish Audio Launches Free TTS API Optimized for Single H200 GPU (2026-07-30, 2 posts)

Fish Audio has launched its S2.1 Pro voice model as a free Text-to-Speech API, supporting 83 languages. This is enabled by extreme inference optimization, achieving 0.17 RTF and 125 audio tokens per second on a single NVIDIA H200 GPU.