Fish Audio Launches Free TTS API: Squeezing 125 Audio Tokens/sec on a Single H200

rohanpaul_ai · x · 2026-07-30

Fish Audio announced that its most advanced voice model, S2.1 Pro, is now available to developers as a free Text-to-Speech (TTS) API, supporting 83 languages with unlimited usage under a Fair Use policy.

To break the industry norm that high-quality voice models must be paid, Fish Audio detailed its underlying compute optimization: every request runs on a single NVIDIA H200 GPU. By using a custom FP8 software stack and a "pingpong" scheduler to eliminate idle time, they achieve an RTF of 0.17 and a throughput of 125 audio tokens per second.

Related event: Fish Audio Launches Free TTS API Optimized for Single H200 GPU(2 posts)→

Original post →

More from Infra

Infra channel →