Full-bandwidth Transformer paper revised: latent feedback nears 1.5x-token performance
_arohan_ · x · 2026-10-06
Authors including Xi Wang and John Langford released a v2 revision of the Full-bandwidth Transformer paper, with unofficial code available.
Key idea: standard autoregressive transformers have a narrow vertical feedback channel — only the sampled token returns to the bottom of the stack each step, while the top-layer hidden state is discarded. This work adds latent feedback: at each decoding step the previous top-layer hidden state is fused with the sampled token embedding via a gated linear unit and fed back as the next input, letting non-verbalized computation re-enter the stack with a fresh depth budget while preserving the standard architecture, KV cache, and LM objective.
Training: a scheduled multi-pass objective introduces latent feedback late in pretraining and mixes a small fraction of deeper feedback passes for stability.
Results: 1B-parameter models trained on up to 400B tokens show improved validation loss, 5-shot LM evals, math and coding generation, and instruction-tuned performance — matching or approaching standard transformers trained with roughly 1.5x more tokens, at negligible per-token decoding overhead.
The revision adds new discussions on debugging training runs, post-training / KV cache replay, and an alternative fusion gate.
More from Models
- Strata engine boosts RTX 3090 prefill 17x, reigniting the local LLM debate — Iory1998 · 2026-10-06
- Reflection unveils Beam: 501B-param open MoE model with 23B active, trained on 23.8T tokens — bigblueboo · 2026-10-06
- Apodex 1.1 Mini hits #4 on OpenRouter trending, 79.6B tokens processed in 4 days — SimonShaoleiDu · 2026-10-06
- "OCR is superhuman now" — a short take on document AI's leap — generativist · 2026-10-06
- Dev returns to Codex after a week on Opus 5.5: "basically unusable" — jacob_posel · 2026-10-06
- CMDB-1500: open-source multimodal benchmark with 1,500 decision-making tasks — Powerful_Buy_4616 · 2026-10-06