Fish Audio Details Inference Stack: 0.17 RTF on a Single H200 GPU

rohanpaul_ai · x · 2026-07-30

Fish Audio revealed the underlying logic of how it can afford to offer its voice model for free: extreme inference optimization.

Hardware & Stack: Every request runs on a single NVIDIA H200 GPU using their custom software stack (fish-scales-ops for fast FP8 math and a 'pingpong' scheduler to prevent idle time).

Metrics: This setup achieves a Real-Time Factor (RTF) of 0.17, meaning it generates audio about 6x faster than real time, outputting 125 audio tokens per second with first sound in 70ms. This high single-GPU throughput drastically reduces the cost per request, making the free tier commercially viable.

Related event: Fish Audio Launches Free TTS API Optimized for Single H200 GPU(2 posts)→

Original post →

More from Infra

Infra channel →