New .wave engine serves 4,800 concurrent ASR streams on one H100 at $0.00045/minute

Ok_boss_labrunz · reddit · 2026-10-04

The .wave team built a persistent-kernel inference engine to serve NVIDIA Nemotron 3.5 ASR Streaming 0.6B at $0.00045/minute.

Core technique — Wave Persistent Kernel (WPK):

Measured on a single H100 SXM 80GB:

The hosted API covers 32 languages with OpenAI Realtime-compatible and Deepgram-compatible endpoints that plug into existing LiveKit and Pipecat setups. Advice for optimizing your own stack: find GPU-host gaps, batch repeated cross-stream work, and watch slow-stream cadence under load.

Original post →

More from Infra

Infra channel →