Cross-layer KV Cache Sharing Research Runs Parallel to Looped Models Done Right II

BlackHC · x · 2026-10-06

The author describes research parallel to Huang et al.'s Looped Models Done Right, Part II, which studies terminal KV reuse (one cache per physical layer), learned training-depth priors, and orthogonal injection. Their own approach shares the KV cache across core blocks for lower memory use and studies donor consistency. The author says they are compute-constrained and invites collaboration, linking the paper.

Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→

Original post →

More from Infra

Infra channel →