Kimi K3 Kernel Achieves 2.05x Speedup Over Official FlashKDA

ChengleiSi · x · 2026-08-19

A technical breakdown reveals that a custom kernel written for Kimi K3's KDA attention achieves a 2.05x speedup over the official FlashKDA on B200 hardware. Now merged into sglang, this kernel features 23 shape-specialized dispatch options and is the best-performing open-source implementation for this workload. The analysis dives into flashinfer-ai/flashinfer to explain how these kernels are produced, consisting of thousands of lines of straightline frozen CUDA and inline PTX.

Original post →

More from Infra

Infra channel →