Tied architectures cost less loss at lower token budgets on fresh data

BlackHC · x · 2026-10-06

Prior work finds gains from looping when data is reused. Using a fresh-data ladder, the author shows a smaller loss cost of parameter tying at lower token budgets, raising the open question of whether shared-cache models can retain those multi-epoch gains.

Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→

Original post →

More from Research

Research channel →