Cohere's CUDA Megakernel Serving Hits 292 tok/s at Batch 1 on a 30B Model

dl_weekly · x · 2026-09-17

Cohere published a deep engineering post on its North Mini Code megakernel serving engine: a 30B model served through one persistent CUDA megakernel reaches 292 tokens/sec at batch 1 — 62% of theoretical peak, versus vLLM's 39%. The approach fuses the entire decode path into a single resident kernel to cut launch and memory-roundtrip overhead, with the biggest wins for low-latency single-request workloads like coding assistants. A rare first-hand serving-optimization writeup.

Original post →

More from Infra

Infra channel →