New .wave engine serves 4,800 concurrent ASR streams on one H100 at $0.00045/minute
Ok_boss_labrunz · reddit · 2026-10-04
The .wave team built a persistent-kernel inference engine to serve NVIDIA Nemotron 3.5 ASR Streaming 0.6B at $0.00045/minute.
Core technique — Wave Persistent Kernel (WPK):
- GPU-resident program: workers execute compiled phases without handing control back to the CPU between ops;
- Cohort streaming: compatible streams share weight reads via weight tiles;
- Resident streaming state: per-stream caches with pre-planned memory layouts cut dynamic allocation.
Measured on a single H100 SXM 80GB:
- Up to 4,800 concurrent streams, 80ms chunks, worst-stream p99 interval 106.3ms (under a 120ms bound);
- Median release-to-completion latency 207–211ms including an 80ms jitter buffer;
- 20× the capacity of NVIDIA's published 240-stream config (not a matched NIM run);
- 0.23% average transcript disagreement vs an fp32 reference.
The hosted API covers 32 languages with OpenAI Realtime-compatible and Deepgram-compatible endpoints that plug into existing LiveKit and Pipecat setups. Advice for optimizing your own stack: find GPU-host gaps, batch repeated cross-stream work, and watch slow-stream cadence under load.
More from Infra
- PrAIvy: A P2P Network to Share Local Ollama Instances Across Users — Matty_za33 · 2026-10-04
- Apple Lisa 1983 vs DGX Spark 2026: the $10k family computer tradition — Thionne_WTZ · 2026-10-04
- AMD GPUs clear over 90% of vLLM gating tests, eroding a key piece of NVIDIA's CUDA moat — AnushElangovan · 2026-10-04
- Go1 Box Targets Compliance-Sensitive Enterprise AI With 50ms Claims and 8,000 Concurrent Requests — swagonflyyyy · 2026-10-04
- CrowdGPT: An Open-Source LLM Trained Without Any Datacenters — Vxtzq1 · 2026-10-04
- Neo4j engineer builds a wearable AI agent on a Raspberry Pi with graph memory — AI Engineer · 2026-10-04