Rubin uses 2:4 attention sparsity to cut bandwidth and storage
nrehiew_ · x · 2026-07-22
The post highlights an attention sparsity idea that keeps the two largest-magnitude values out of every local group of four, stores their indices, and zeros the rest.
It says Rubin can then run MMA directly on the compressed tensor plus indices, reducing bandwidth and storage overhead. The attached diagram shows the compression format and how the sparse feature is inserted into transformer attention and MLP blocks.
More from Infra
- Stress-Testing an AI Gateway With 100 Concurrent Agents: All HTTP 200, Continuity Collapsed — its_vayishu · 2026-09-12
- llama.cpp PR adds missing AMD GCN MMQ config, boosting MI50/MI60 inference — pmttyji · 2026-09-12
- TensorSharp hits 41 tok/s decoding DeepSeek V4.1 Flash on 8× A40 — fuzhongkai · 2026-09-12
- What actually runs AI models at the edge in 2026: Mac mini, DGX Spark, iPhone 17 Pro — MaziyarPanahi · 2026-09-12
- Running Qwen3.8 Flash Next on dual RTX 3090: full llama.cpp config shared for tuning — ChopSticksPlease · 2026-09-12
- UAE redesigns 5GW AI campus with bunkers and air defenses after Iranian strikes on Gulf cloud facilities — mark_k · 2026-09-12