DeepSeek V4.1 Flash shows massive kernel-engineering gains, hits 4th on KernelBench-CUDA

teortaxesTex · x · 2026-09-12

DeepSeek V4.1 Flash reportedly shows a massive leap in kernel engineering, far ahead of GLM-5.3 Flash. In a KernelBench-CUDA test it wrote a Native Sparse Attention kernel for RTX PRO 6000 reaching 0.50 of dense-equivalent roofline (4th place; Fable 5.1: 1.06, Opus 5: 1.04), using inline PTX with mma.sync bf16, ldmatrix, xor-swizzled cp.async, and a fused fp32 top-8 block-scoring prologue. The author notes it executes 78% of the causal block triangle where semantics need only 14% at 8K context, which explains the gap to the top three.

Related event: DeepSeek 4.1 Flash Shows Big Gains in Kernel Engineering and Cost(2 posts)→

Original post →

More from Models

Models channel →