Rubin uses 2:4 attention sparsity to cut bandwidth and storage

nrehiew_ · x · 2026-07-22

The post highlights an attention sparsity idea that keeps the two largest-magnitude values out of every local group of four, stores their indices, and zeros the rest.

It says Rubin can then run MMA directly on the compressed tensor plus indices, reducing bandwidth and storage overhead. The attached diagram shows the compression format and how the sparse feature is inserted into transformer attention and MLP blocks.

Original post →

More from Infra

Infra channel →