WhiteMatter: All-to-All Cross-Layer KV Sharing Matches Bigger Models With Half the Cache
INK-USC · hf · 2026-09-30
- Motivation: In Transformers, each layer can only access past-token representations from the same depth, limiting reuse of computed information.
- Method: WhiteMatter lets every layer draw representations from any depth via a learned mixer, combining them into shared KV cache channels across layers, which also shrinks cache size.
- Results: With the same training tokens, full-cache WhiteMatter matches a standard Transformer with 50% more layers; with half the KV cache it beats matched standard Transformers up to 1.3B parameters.
- Engineering: Cross-layer dependencies slow training/prefill; cyclic iteration (interleaved token groups updated in turn) converges 12.5x faster than standard Jacobi iteration.
More from Infra
- Nvidia CEO Jensen Huang: Data centers are now 'superintelligence factories' — Polymarket · 2026-09-30
- Ollama 0.35 ships with Nimble support for fast local Mac inference — Technovangelist · 2026-09-30
- AMD's Hyperloom and ROCm 10: AI agents tune GPU kernels overnight with accuracy checks — AnushElangovan · 2026-09-30
- Efficient raises $97M Series B to rethink computing from physical AI to data centers — Sethwinterroth · 2026-09-30
- Developer: OpenAI is taking compute from paying users for consumer agents — arthurcolle · 2026-09-30
- Prime Intellect to deploy on NVIDIA's new Vera CPU in first wave — eliebakouch · 2026-09-30