Tied transformer cuts params 228.5M to 136.2M for +0.044 nats; late-training KV match removes 81-89% penalty
BlackHC · x · 2026-10-06
A former DeepMind researcher's first post-departure, compute-constrained project: an LSTM-style tied transformer with one shared core KV cache, studying how to train for cache histories produced recursively at inference.
- At width 1024, tying cuts parameters from 228.5M to 136.2M at similar forward FLOPs (+0.0438 nats); tied-1024 beats untied-768 at 136M params using 1.66x reference FLOPs (one seed).
- Matching the two passes' donor keys/values in the final 10% of training removes 81-89% of the continuation penalty at widths 512/768/1024, with full-cache loss rising only 0.0011-0.0025 nats.
- Works without LSTM-style carry too: donor consistency improves cold sequential loss by -0.1374/-0.0814/-0.0271 nats for gated carry, plain recurrence, and input injection (one seed each).
- Negative result: a state-cosine adaptive-depth stopping rule picks 4.75 passes at loss 3.3816, worse than fixed depth 4 at 3.3761.
- Fresh-data ladder shows a smaller loss cost of tying at lower token budgets; open question: can shared-cache models retain multi-epoch gains?
Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→
More from Infra
- Cloudflare Lets Workers Connect to Artifacts Repos, Cutting GitHub Out of the Build Pipeline — threepointone · 2026-10-06
- Vultr books $1.2B AMD AI rack order as buyers reserve capacity years ahead — shashib · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- NanoGPT speedrun sets record: 11.3% faster via architecture-only change, paper coming — yoavartzi · 2026-10-06
- NVIDIA's CANTO Predicts Aerodynamics Directly From CAD, Cuts Pressure Error 20% — JeanKossaifi · 2026-10-06
- Epoch AI: compute could soon support hundreds of millions to billions of AI agents — Jsevillamol · 2026-10-06