Real-time AI serving costs up to 56x more than needed — one engine hits 56 sessions per H100

Ok_boss_labrunz · reddit · 2026-10-02

DotWave published two technical write-ups showing that full-duplex voice models — which must emit a frame every 80ms even during silence — break the request-based assumptions of vLLM, SGLang, and TensorRT-LLM. NVIDIA's reference stack leaves the GPU idle 75% of the time and drops 8–15% of audio frames at just two concurrent sessions, so each call effectively pays for a whole H100. Their fix keeps the model as a resident GPU program, batches all sessions due on the same tick, and fixes scheduling at compile time. Same weights and precision, but 56 concurrent sessions per H100 (vs 1), 147.4ms p99 per 160ms beat, zero missed deadlines across 84,000 session-beats — roughly 98% less GPU per conversation.

Original post →

More from Infra

Infra channel →