Cross-layer KV Cache Sharing Research Runs Parallel to Looped Models Done Right II
BlackHC · x · 2026-10-06
The author describes research parallel to Huang et al.'s Looped Models Done Right, Part II, which studies terminal KV reuse (one cache per physical layer), learned training-depth priors, and orthogonal injection. Their own approach shares the KV cache across core blocks for lower memory use and studies donor consistency. The author says they are compute-constrained and invites collaboration, linking the paper.
Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→
More from Infra
- Pi-hole-class DNS ad-blocker runs on a $2 ESP32-C3 with 537k domains in flash — M-Abozaid · 2026-10-06
- Cloudflare Lets Workers Connect to Artifacts Repos, Cutting GitHub Out of the Build Pipeline — threepointone · 2026-10-06
- Vultr books $1.2B AMD AI rack order as buyers reserve capacity years ahead — shashib · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- NanoGPT speedrun sets record: 11.3% faster via architecture-only change, paper coming — yoavartzi · 2026-10-06
- NVIDIA's CANTO Predicts Aerodynamics Directly From CAD, Cuts Pressure Error 20% — JeanKossaifi · 2026-10-06