Gated DeltaNet-2 gets full cuDNN support, ~3x faster end-to-end training on NVIDIA GPUs

ZGojcic · x · 2026-08-24

NVIDIA's cuDNN team shipped full support for Gated DeltaNet-2 (GDN-2), covering not just prefill but a highly optimized backward pass built on CUTLASS primitives, released as part of CUTLASS 4.7.0.

On GB300 (BF16, batch 4, 64 heads, d=128) vs. the FLA Triton implementation: forward up to 6.4x faster, backward up to 2.8x, and 3x faster end-to-end training. Throughput stays flat from 2K to 32K sequence lengths, meaning the kernels are compute-bound where they should be.

The accompanying paper, Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention, separates the scalar tie between erasing and writing in delta-rule linear attention via a channel-wise erase gate and write gate, generalizing both Gated DeltaNet and Kimi Delta Attention (KDA). At 1.3B params / 100B FineWeb-Edu tokens it beats Mamba-2, GDN, KDA and Mamba-3 variants overall on language modeling.

Original post →

More from Infra

Infra channel →