Fish Audio Launches Free TTS API: Squeezing 125 Audio Tokens/sec on a Single H200
rohanpaul_ai · x · 2026-07-30
Fish Audio announced that its most advanced voice model, S2.1 Pro, is now available to developers as a free Text-to-Speech (TTS) API, supporting 83 languages with unlimited usage under a Fair Use policy.
To break the industry norm that high-quality voice models must be paid, Fish Audio detailed its underlying compute optimization: every request runs on a single NVIDIA H200 GPU. By using a custom FP8 software stack and a "pingpong" scheduler to eliminate idle time, they achieve an RTF of 0.17 and a throughput of 125 audio tokens per second.
Related event: Fish Audio Launches Free TTS API Optimized for Single H200 GPU(2 posts)→
More from Infra
- World's Largest Physical AI Data Engine Cranks Up Capacity — pduan · 2026-07-31
- SGLang Partners with Google and RadixArk to Deliver Native TPU Inference for LLMs — ying11231 · 2026-07-31
- LLM Routers Emerge as a Distinct Service Category for Cost and Performance — CackleRooster · 2026-07-31
- Together AI Webinar: Deploying Open-Weight Models in Production — togethercompute · 2026-07-31
- NVIDIA Launches Jetson AGX Thor: The Ultimate Robotics Brain for Humanoids — NVIDIA Developer · 2026-07-31
- AI Data Center Noise Hits 105 Decibels, Sparking Lawsuits Against Tech Giants — mkheck · 2026-07-31