Tied transformer cuts params 228.5M to 136.2M for +0.044 nats; late-training KV match removes 81-89% penalty

BlackHC · x · 2026-10-06

A former DeepMind researcher's first post-departure, compute-constrained project: an LSTM-style tied transformer with one shared core KV cache, studying how to train for cache histories produced recursively at inference.

Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→

Original post →

More from Infra

Infra channel →