Ex-DeepMind researcher demos tied transformer with a shared core KV cache
BlackHC · x · 2026-10-06
Former Google DeepMind researcher BlackHC shares his first post-departure research project: an LSTM-style tied transformer with one shared core KV cache.
The central question is how to train for the cache histories the model recursively produces at inference. It builds on LCKV (final-layer caches + KV matching), CLA (cross-layer reuse), and Huginn (recurrent depth), but specifically tests cache-history mismatch and repair in a tied core using LSTM cells. He's compute-constrained and open to collaboration.
Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→
More from Infra
- Pi-hole-class DNS ad-blocker runs on a $2 ESP32-C3 with 537k domains in flash — M-Abozaid · 2026-10-06
- Cloudflare Lets Workers Connect to Artifacts Repos, Cutting GitHub Out of the Build Pipeline — threepointone · 2026-10-06
- Vultr books $1.2B AMD AI rack order as buyers reserve capacity years ahead — shashib · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- NanoGPT speedrun sets record: 11.3% faster via architecture-only change, paper coming — yoavartzi · 2026-10-06
- NVIDIA's CANTO Predicts Aerodynamics Directly From CAD, Cuts Pressure Error 20% — JeanKossaifi · 2026-10-06