Kimi K3 Kernel Achieves 2.05x Speedup Over Official FlashKDA
ChengleiSi · x · 2026-08-19
A technical breakdown reveals that a custom kernel written for Kimi K3's KDA attention achieves a 2.05x speedup over the official FlashKDA on B200 hardware. Now merged into sglang, this kernel features 23 shape-specialized dispatch options and is the best-performing open-source implementation for this workload. The analysis dives into flashinfer-ai/flashinfer to explain how these kernels are produced, consisting of thousands of lines of straightline frozen CUDA and inline PTX.
More from Infra
- MCDMA Enables Direct RDMA Between NVIDIA Spark and Apple Silicon — StephanSturges · 2026-08-19
- Cursor's deep dive on Git at any scale: running Origin like a database — vmg · 2026-08-19
- Gemma 4 Powers Local Enterprise RFP Agent — victormustar · 2026-08-19
- Vector database Qdrant surpasses 250M downloads, recognized by Forrester — qdrant_engine · 2026-08-19
- Mojo programming language is now open source, combining Python ease with C++ performance — mgill25 · 2026-08-19
- Infra Behind Krea 2: Tensor Cores and Crash Philosophy — AI Engineer · 2026-08-19