Open-source KDA-B200 kernel claims up to 1.59× speedup on NVIDIA Blackwell GPUs
xennygrimmato_ · x · 2026-07-28
INT21-AI has open-sourced KDA-B200, a forward-only Kimi Delta Attention implementation for NVIDIA Blackwell GPUs.
- It is written in plain CUDA C++ and PTX, with no CUTLASS or CUTE runtime/build dependency.
- The repo reports a 1.42× geometric-mean speedup over the CUTLASS-based FlashKDA path across six same-host scenarios, and up to 1.59× in the cited benchmark.
- The implementation was validated on an NVIDIA B200 and targets sm100a hardware.
The post is essentially a kernel-performance announcement: a hand-rolled GPU path for KDA that improves throughput on Blackwell-class hardware.
Related event: Open-Source PTX KDA Kernel Released for Blackwell(2 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23