An LSTM-style tied transformer with one shared core KV cache

BlackHC · x · 2026-10-06

A former Google DeepMind researcher's compact research project: an LSTM-style tied transformer where the whole model shares a single core KV cache. Built on LCKV (final-layer caches + KV matching), CLA (cross-layer reuse), and Huginn (recurrent depth), it tests cache-history mismatch and repair in a tied core using LSTM cells, tackling how to train for cache histories the model produces recursively at inference.

Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→

Original post →

More from Infra

Infra channel →