Cohere's megakernel serving engine hits 1.58x vLLM on a single H100
cohere · x · 2026-09-09
- Cohere announced its North Mini Code Megakernel serving engine with a blog post explaining what a megakernel is and how the system was built.
- Reported results: with North Mini Code (BF16 on a single H100), the engine achieves 1.58x vs vLLM at batch size 1, and 1.25x–1.41x end-to-end serving throughput at BS=8.
- The work targets inference-stack optimization for agentic coding model deployment.
Related event: Cohere open-sources Megakernel inference engine, 1.58x faster than vLLM(2 posts)→
More from Infra
- NVIDIA ships CUDA Python 1.0 with stable APIs, making Python first-class for CUDA — PyTorch · 2026-09-09
- Inference is turning GPU compute into a tradable commodity — ArtificialAnlys · 2026-09-09
- Cohere open-sources megakernel serving engine, up to 1.58x faster than vLLM — cohere · 2026-09-09
- Solving Navier-Stokes cost 130B output tokens — up to $18M depending on model pricing — mitsuhiko · 2026-09-09
- Baseten cuts delta weight syncs for frontier models to under 40 seconds — baseten · 2026-09-09
- Dell says DRAM, NAND shortages persist and nearly all leading-node products are constrained — Beth_Kindig · 2026-09-09