Cohere open-sources Megakernel inference engine, 1.58x faster than vLLM
Cohere open-sourced North Mini Code, a complete LLM inference system built around a decode megakernel that runs up to 1.58x faster than vLLM on a single H100.
2026-09-09 ~ 2026-09-09 · 2 related posts
- Cohere open-sources megakernel serving engine, up to 1.58x faster than vLLM — cohere · 2026-09-09
- Cohere's megakernel serving engine hits 1.58x vLLM on a single H100 — cohere · 2026-09-09