Tied architectures cost less loss at lower token budgets on fresh data
BlackHC · x · 2026-10-06
Prior work finds gains from looping when data is reused. Using a fresh-data ladder, the author shows a smaller loss cost of parameter tying at lower token budgets, raising the open question of whether shared-cache models can retain those multi-epoch gains.
Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→
More from Research
- Studies: humans deny AI consciousness even with identical behavior; AI vision misses illusions primates catch — MacrinePhD · 2026-10-06
- SFT then RL doesn't fix agent looping: 29% of runs hit turn cap vs 0% for RL alone — VikParuchuri · 2026-10-06
- RL Post-Training Eliminates Agent Tool-Call Loops: 92% Loop Rate Drops to 0 — VikParuchuri · 2026-10-06
- Watch, Infer, Coordinate: robots infer a partner's physical limits from watching teamwork, then coordinate zero-shot — mangahomanga · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- Math lacks empirical tradition: Wolfskehl Prize drew 1,000 wrong Fermat proofs — RexDouglass · 2026-10-06