Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache
BlackHC, a former Google DeepMind researcher, has released his first research project since leaving (self-described as a small, compute-limited effort): an LSTM-style tied Transformer that keeps only a single shared core KV cache for the whole model, with the core question being how to train against the cache history the model recursively produces at inference time.
Confirmed
- The work builds on three existing lines: LCKV (last-layer cache + KV matching), CLA (cross-layer reuse), and Huginn (recurrent depth), using an LSTM cell within the tied core to test cache-history mismatch and fixes for it.
- Key numbers: tying saves roughly 40% of parameters at a cost of only +0.04 nats of loss; applying KV matching during the final 10% of training eliminates about 80% of the train–inference mismatch penalty.
- The research is parallel and independent of Huang et al.'s "Looped Models Done Right, Part II"; both aim to reduce inference-time KV cache memory footprint but take different technical routes, and BlackHC has shared the paper link.
Why it matters
- KV cache memory footprint is a major bottleneck in current inference costs, and cross-layer/recurrent KV cache reuse is an active direction for cutting inference overhead; this work provides empirical data on tied weights plus recursive-cache training.
- The author notes that even small, compute-limited experiments can quantify the sources and repairability of cache-history mismatch, offering a methodological reference for larger-scale validation.
2026-10-06 ~ 2026-10-06 · 8 related posts
Primary sources
- Ex-DeepMind researcher demos tied transformer with a shared core KV cache — BlackHC ·
- Paper out: tied transformer with shared core KV cache, building on LCKV, CLA, Huginn — BlackHC ·
- Tied transformer cuts params 228.5M to 136.2M for +0.044 nats; late-training KV match removes 81-89% penalty — BlackHC ·
- [source] Ex-DeepMind researcher demos tied transformer with a shared core KV cache — BlackHC · 2026-10-06
- [source] Tied transformer cuts params 228.5M to 136.2M for +0.044 nats; late-training KV match removes 81-89% penalty — BlackHC · 2026-10-06
- Donor-consistency repair helps without LSTM-style carry across three variants — BlackHC · 2026-10-06
- Tied architectures cost less loss at lower token budgets on fresh data — BlackHC · 2026-10-06
- Negative result: state-cosine adaptive depth stopping rule hurts both loss and passes — BlackHC · 2026-10-06
- An LSTM-style tied transformer with one shared core KV cache — BlackHC · 2026-10-06
- [source] Paper out: tied transformer with shared core KV cache, building on LCKV, CLA, Huginn — BlackHC · 2026-10-06
- Cross-layer KV Cache Sharing Research Runs Parallel to Looped Models Done Right II — BlackHC · 2026-10-06