Cohere open-sources megakernel serving engine, up to 1.58x faster than vLLM

cohere · x · 2026-09-09

Cohere has released the first fully-fledged LLM serving system built around a decode megakernel, fully open-source on GitHub.

Related event: Cohere open-sources Megakernel inference engine, 1.58x faster than vLLM(2 posts)→

Original post →

More from Infra

Infra channel →