Swapping matmul for associative-algebra layers boosts 110M LM throughput 7.8%
Ilya Koziev · hf · 2026-09-29
Instead of fixing the product and searching for cheaper evaluation like fast matmul algorithms, this work lets a Transformer's projections use a different, cheaper product altogether: an associative-algebra construction replacing dense matrix multiplication with a sparser interaction table over the same weight blocks.
Key points:
- The construction has quadratic arithmetic in matrix dimension with fixed block size, is provably optimal in bilinear rank via the Alder–Strassen bound, and works with causal masking and KV-cached decoding;
- Empirical test: two 110M decoder-only LMs trained on the same recipe and 12.3B-token budget, differing only in feed-forward layers — the algebraic model gains 6.2–7.8% end-to-end generation throughput at slightly lower downstream scores;
- Framed as a feasibility and trainability check at small scale.
More from Infra
- Lambda's H100s keep stocking out as reserved clusters get priority, pushing European teams elsewhere — YvesMulkers · 2026-09-29
- 23 of 128 Bittensor subnets now make real money; top GPU subnet billed $964K last month — bittingthembits · 2026-09-29
- Firebase SDK is crashing large numbers of iOS apps since this morning — pranshuchittora · 2026-09-29
- Shaw mocks data center opponents: hating compute while using the internet is incoherent — zealcaiden · 2026-09-29
- Modeling 1B agent VMs by 2030: what personal AI agents mean for CPU demand — AccBalanced · 2026-09-29
- Tessera: retrieval-driven KV cache reuse cuts RAG serving TTFT by up to 3.6x — _reachsumit · 2026-09-29