Cohere's CUDA Megakernel Serving Hits 292 tok/s at Batch 1 on a 30B Model
dl_weekly · x · 2026-09-17
Cohere published a deep engineering post on its North Mini Code megakernel serving engine: a 30B model served through one persistent CUDA megakernel reaches 292 tokens/sec at batch 1 — 62% of theoretical peak, versus vLLM's 39%. The approach fuses the entire decode path into a single resident kernel to cut launch and memory-roundtrip overhead, with the biggest wins for low-latency single-request workloads like coding assistants. A rare first-hand serving-optimization writeup.
More from Infra
- Data centers now supply roughly 45% of local tax revenue in Loudoun County, Virginia — Polymarket · 2026-09-17
- Keeping vLLM's prefix cache warm between agent turns — bolts98 · 2026-09-17
- Running Qwen3.8-Flash-Next on 2x5080: should I move from llama.cpp to vLLM? — whatyathinkk · 2026-09-17
- Nvidia's Jensen Huang says chip sales will double next year — zephyr_z9 · 2026-09-17
- GlobalFoundries and Marvell expand Vermont SiGe capacity for AI optical networking — zephyr_z9 · 2026-09-17
- $26,100 desktop AI datacenter: dual RTX PRO 6000 Blackwell workstation goes open source — dee_hw · 2026-09-17