Rubin uses 2:4 attention sparsity to cut bandwidth and storage
nrehiew_ · x · 2026-07-22
The post highlights an attention sparsity idea that keeps the two largest-magnitude values out of every local group of four, stores their indices, and zeros the rest.
It says Rubin can then run MMA directly on the compressed tensor plus indices, reducing bandwidth and storage overhead. The attached diagram shows the compression format and how the sparse feature is inserted into transformer attention and MLP blocks.
More from Infra
- Polymarket puts 77% odds on a U.S. state data center moratorium this year — Polymarket · 2026-07-23
- Framework previews a 192GB Ryzen AI desktop for running large models locally — AnushElangovan · 2026-07-23
- CoreWeave says Vera Rubin NVL72 delivers 10x better tokens per megawatt — mark_k · 2026-07-23
- Kimi says MI355X serving changes cut p99 TTFT by 3.2× and lift throughput 7.7% — AnushElangovan · 2026-07-23
- TargonOS launches confidential GPU and CPU VMs with hardware-attested isolation — JosephJacks_ · 2026-07-23
- Cloud infrastructure is splintering into multi-cloud, hybrid and sovereign layers — BenBajarin · 2026-07-23