Gated DeltaNet-2 gets full cuDNN support, ~3x faster end-to-end training on NVIDIA GPUs
ZGojcic · x · 2026-08-24
NVIDIA's cuDNN team shipped full support for Gated DeltaNet-2 (GDN-2), covering not just prefill but a highly optimized backward pass built on CUTLASS primitives, released as part of CUTLASS 4.7.0.
On GB300 (BF16, batch 4, 64 heads, d=128) vs. the FLA Triton implementation: forward up to 6.4x faster, backward up to 2.8x, and 3x faster end-to-end training. Throughput stays flat from 2K to 32K sequence lengths, meaning the kernels are compute-bound where they should be.
The accompanying paper, Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention, separates the scalar tie between erasing and writing in delta-rule linear attention via a channel-wise erase gate and write gate, generalizing both Gated DeltaNet and Kimi Delta Attention (KDA). At 1.3B params / 100B FineWeb-Edu tokens it beats Mamba-2, GDN, KDA and Mamba-3 variants overall on language modeling.
More from Infra
- AMD Instinct MI210 for Local LLMs: A Cost-Effective Choice? — OvertaxedOne · 2026-08-24
- Wells Fargo sees Broadcom AI chip revenue at $205B by FY28, far above consensus — Beth_Kindig · 2026-08-24
- Disaggregated LPDDR memory: A new path for AI infrastructure — jwt0625 · 2026-08-24
- Study notes: how speculative decoding accelerates LLM inference without quality loss — helloiamleonie · 2026-08-24
- Global AI Power Demand Projected to Surge 1,100% by 2033 — KyeGomezB · 2026-08-24
- RTX 3080 Memory OC to +1200 Boosts Flux Generation Efficiency — MakionGarvinus · 2026-08-24