Swapping matmul for associative-algebra layers boosts 110M LM throughput 7.8%

Ilya Koziev · hf · 2026-09-29

Instead of fixing the product and searching for cheaper evaluation like fast matmul algorithms, this work lets a Transformer's projections use a different, cheaper product altogether: an associative-algebra construction replacing dense matrix multiplication with a sparser interaction table over the same weight blocks.

Key points:

Original post →

More from Infra

Infra channel →