Moonshot opensources FlashKDA, claiming 1.72×–2.22× H20 prefill gains
ctjlewis · x · 2026-07-27
Moonshot AI open-sourced FlashKDA, a high-performance CUTLASS-based implementation of Kimi Delta Attention kernels.
- The repo is designed as a drop-in backend for flash-linear-attention.
- The team says it delivers 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20.
- Requirements listed in the repo include SM90+, CUDA 12.9+, and PyTorch 2.4+.
More from Infra
- Kimi K3 launches on SGLang with 423 tok/s and 11 cloud partners — ying11231 · 2026-07-28
- Claude chat indexing incident sparks a push for confidential-compute AI — bittingthembits · 2026-07-27
- Moonshot open-sources MoonEP as open models vs closed labs debate intensifies — KyeGomezB · 2026-07-27
- Kimi K3 goes live on Nebius with 1M-token context and a 57 AA score — teortaxesTex · 2026-07-27
- AI agent finds a longstanding Bun Node-compat bug in `child_process.spawn` — steipete · 2026-07-27
- llama.cpp adds support for Nanbeige4.2 in pull request 25994 — pmttyji · 2026-07-27