DeepSeek V4.1 Flash shows massive kernel-engineering gains, hits 4th on KernelBench-CUDA
teortaxesTex · x · 2026-09-12
DeepSeek V4.1 Flash reportedly shows a massive leap in kernel engineering, far ahead of GLM-5.3 Flash. In a KernelBench-CUDA test it wrote a Native Sparse Attention kernel for RTX PRO 6000 reaching 0.50 of dense-equivalent roofline (4th place; Fable 5.1: 1.06, Opus 5: 1.04), using inline PTX with mma.sync bf16, ldmatrix, xor-swizzled cp.async, and a fused fp32 top-8 block-scoring prologue. The author notes it executes 78% of the causal block triangle where semantics need only 14% at 8K context, which explains the gap to the top three.
Related event: DeepSeek 4.1 Flash Shows Big Gains in Kernel Engineering and Cost(2 posts)→
More from Models
- My GPT live voice agent actually argued with me — first real pushback — evielync · 2026-09-12
- kalomaze: Poor model transfer reflects narrow human problem selection, not failed generalization — kalomaze · 2026-09-12
- Sakana AI Ships Fugu Ultra v2, Multi-Model Orchestration Returns to World-Class Performance — SakanaAILabs · 2026-09-12
- Grok fails to find a user's own tweet after multiple tries; ChatGPT nails it instantly — SimonBalmain · 2026-09-12
- HSVSphere slams Opus 5: 'it actually makes you lose time' — yacineMTB · 2026-09-12
- GPT 6 Astra demand overwhelms OpenAI, anti-AI backlash has 'completely failed' — Tolopono · 2026-09-12