An LSTM-style tied transformer with one shared core KV cache
BlackHC · x · 2026-10-06
A former Google DeepMind researcher's compact research project: an LSTM-style tied transformer where the whole model shares a single core KV cache. Built on LCKV (final-layer caches + KV matching), CLA (cross-layer reuse), and Huginn (recurrent depth), it tests cache-history mismatch and repair in a tied core using LSTM cells, tackling how to train for cache histories the model produces recursively at inference.
Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→
More from Infra
- Pi-hole-class DNS ad-blocker runs on a $2 ESP32-C3 with 537k domains in flash — M-Abozaid · 2026-10-06
- Cloudflare Lets Workers Connect to Artifacts Repos, Cutting GitHub Out of the Build Pipeline — threepointone · 2026-10-06
- Vultr books $1.2B AMD AI rack order as buyers reserve capacity years ahead — shashib · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- NanoGPT speedrun sets record: 11.3% faster via architecture-only change, paper coming — yoavartzi · 2026-10-06
- NVIDIA's CANTO Predicts Aerodynamics Directly From CAD, Cuts Pressure Error 20% — JeanKossaifi · 2026-10-06