Real-time AI serving costs up to 56x more than needed — one engine hits 56 sessions per H100
Ok_boss_labrunz · reddit · 2026-10-02
DotWave published two technical write-ups showing that full-duplex voice models — which must emit a frame every 80ms even during silence — break the request-based assumptions of vLLM, SGLang, and TensorRT-LLM. NVIDIA's reference stack leaves the GPU idle 75% of the time and drops 8–15% of audio frames at just two concurrent sessions, so each call effectively pays for a whole H100. Their fix keeps the model as a resident GPU program, batches all sessions due on the same tick, and fixes scheduling at compile time. Same weights and precision, but 56 concurrent sessions per H100 (vs 1), 147.4ms p99 per 160ms beat, zero missed deadlines across 84,000 session-beats — roughly 98% less GPU per conversation.
More from Infra
- Armada becomes Palantir's first Certified Modular Data Center partner for off-grid sovereign AI — brucemacv · 2026-10-02
- Nvidia hits record $5.7T market cap, adds record $150B buyback on AI agent optimism — kimmonismus · 2026-10-02
- shardr: a content-addressed model store with BitTorrent sync and OpenAI-compatible serving — Cyb3erDudu · 2026-10-02
- Cloudflare adds Analytics SQL binding to query analytics datasets from Workers — ritakozlov · 2026-10-02
- llama.cpp adds decision models: /v1/systemone scores options in a single forward pass — ggerganov · 2026-10-02
- Deutsche Bank Initiates FormFactor at Buy With $200 Target on Nvidia GPU Probe Card Share — demian_ai · 2026-10-02