Cohere's megakernel serving engine hits 1.58x vLLM on a single H100

cohere · x · 2026-09-09

Related event: Cohere open-sources Megakernel inference engine, 1.58x faster than vLLM(2 posts)→

Original post →

More from Infra

Infra channel →