Open-source KDA-B200 kernel claims up to 1.59× speedup on NVIDIA Blackwell GPUs
xennygrimmato_ · x · 2026-07-28
INT21-AI has open-sourced KDA-B200, a forward-only Kimi Delta Attention implementation for NVIDIA Blackwell GPUs.
- It is written in plain CUDA C++ and PTX, with no CUTLASS or CUTE runtime/build dependency.
- The repo reports a 1.42× geometric-mean speedup over the CUTLASS-based FlashKDA path across six same-host scenarios, and up to 1.59× in the cited benchmark.
- The implementation was validated on an NVIDIA B200 and targets sm100a hardware.
The post is essentially a kernel-performance announcement: a hand-rolled GPU path for KDA that improves throughput on Blackwell-class hardware.
Related event: Open-Source PTX KDA Kernel Released for Blackwell(2 posts)→
More from Infra
- Jensen Huang backs open-source LLMs, but GPU costs remain the real barrier — pascalefung · 2026-07-28
- Google’s AlloyDB Omni demo runs fully offline with Gemma and TimesFM — rseroter · 2026-07-28
- Nvidia’s $750 billion AI deals reignite fears of circular financing — brainquantum · 2026-07-28
- SSI Signs $410M Compute Deal with Amazon, Bulk of Its Fundraising — TechCrunch AI · 2026-07-28
- Taiwan Detains Nvidia Employee in Widening AI Server Smuggling Probe — The Decoder · 2026-07-28
- RBC says a DIY AI replacement for Office could cost 11 times more over five years — TiernanRayTech · 2026-07-28