Full Bandwidth Transformer could preserve information in final layers, says Langford
JohnCLangford · x · 2026-09-06
Continuing his thread, John Langford argues a Full Bandwidth Transformer could use early layers for deep compute while final layers preserve relevant information instead of discarding it — yielding better depth utilization and more depth, with gradients free to quench irrelevant prior-token information.
Related event: Langford Explores Information Dropping in Transformer Architectures(3 posts)→
More from Research
- Startup Mostik bridges models in latent space, tops ARC-AGI 3 leaderboard — emilahlback · 2026-09-06
- DeepMind launches WeatherNext 3, a weather model trained on real-time observations — ZoubinGhahrama1 · 2026-09-06
- Slower enzyme, more product: DLCatalysis lifts NeuAc yield from 8.54 to 9.27 g/L — bravo_abad · 2026-09-06
- ORNL shows AI assembling artificial graphene atom-by-atom autonomously for 25+ hours — TinfoilTricorn · 2026-09-06
- "Why I'm building thermodynamic computers" draws AI community attention — TinfoilTricorn · 2026-09-06
- Researcher calls on RL teams to add refactoring and deletion tasks to agent training — kuza55 · 2026-09-06