Full-bandwidth Transformer paper revised: latent feedback nears 1.5x-token performance

_arohan_ · x · 2026-10-06

Authors including Xi Wang and John Langford released a v2 revision of the Full-bandwidth Transformer paper, with unofficial code available.

Key idea: standard autoregressive transformers have a narrow vertical feedback channel — only the sampled token returns to the bottom of the stack each step, while the top-layer hidden state is discarded. This work adds latent feedback: at each decoding step the previous top-layer hidden state is fused with the sampled token embedding via a gated linear unit and fed back as the next input, letting non-verbalized computation re-enter the stack with a fresh depth budget while preserving the standard architecture, KV cache, and LM objective.

Training: a scheduled multi-pass objective introduces latent feedback late in pretraining and mixes a small fraction of deeper feedback passes for stability.

Results: 1B-parameter models trained on up to 400B tokens show improved validation loss, 5-shot LM evals, math and coding generation, and instruction-tuned performance — matching or approaching standard transformers trained with roughly 1.5x more tokens, at negligible per-token decoding overhead.

The revision adds new discussions on debugging training runs, post-training / KV cache replay, and an alternative fusion gate.

Original post →

More from Models

Models channel →