Fish Audio Details Inference Stack: 0.17 RTF on a Single H200 GPU
rohanpaul_ai · x · 2026-07-30
Fish Audio revealed the underlying logic of how it can afford to offer its voice model for free: extreme inference optimization.
Hardware & Stack: Every request runs on a single NVIDIA H200 GPU using their custom software stack (fish-scales-ops for fast FP8 math and a 'pingpong' scheduler to prevent idle time).
Metrics: This setup achieves a Real-Time Factor (RTF) of 0.17, meaning it generates audio about 6x faster than real time, outputting 125 audio tokens per second with first sound in 70ms. This high single-GPU throughput drastically reduces the cost per request, making the free tier commercially viable.
Related event: Fish Audio Launches Free TTS API Optimized for Single H200 GPU(2 posts)→
More from Infra
- AI Demand Far Outpaces Supply as Infrastructure Becomes the Bottleneck — NinaDSchick · 2026-07-30
- Open-source project runs 120B MoE models on smartphones at 6 tokens/s — dai_app · 2026-07-30
- How Open-Source Small Models Power Edge AI for Wildfire Detection in France — import_jmr · 2026-07-30
- Post-training pain points: messy parallel experiments, dependency hell, high GPU costs — Pitiful-Minute-2818 · 2026-07-30
- Nscale Acquires Anyscale to Own More of the AI Compute Stack — TechCrunch AI · 2026-07-30
- Troubleshooting RAM/VRAM Allocation for MTP in llama.cpp — xornullvoid · 2026-07-30