Cohere open-sources Megakernel inference engine, 1.58x faster than vLLM

Cohere open-sourced North Mini Code, a complete LLM inference system built around a decode megakernel that runs up to 1.58x faster than vLLM on a single H100.

2026-09-09 ~ 2026-09-09 · 2 related posts