Ex-DeepMind researcher demos tied transformer with a shared core KV cache

BlackHC · x · 2026-10-06

Former Google DeepMind researcher BlackHC shares his first post-departure research project: an LSTM-style tied transformer with one shared core KV cache.

The central question is how to train for the cache histories the model recursively produces at inference. It builds on LCKV (final-layer caches + KV matching), CLA (cross-layer reuse), and Huginn (recurrent depth), but specifically tests cache-history mismatch and repair in a tied core using LSTM cells. He's compute-constrained and open to collaboration.

Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→

Original post →

More from Infra

Infra channel →