Langford details Full Bandwidth Transformer and addresses OpenAI Astra rumor
Machine learning veteran John Langford, associated with Vowpal Wabbit and MLic, posted a series of updates on September 4 introducing his team's progress and technical thinking on the Full Bandwidth Transformer (FBT) while also addressing rumors that "OpenAI's new project Astra uses FBT." Two core takeaways: FBT unexpectedly turned out not to need nextlat in implementation, and someone from OpenAI had come asking where the idea came from, which Langford found odd.
Confirmed
- Langford said that while developing FBT there were two surprises: nextlat was not needed, and training went remarkably smoothly.
- On dropping nextlat: the nextlat paper's proof implicitly converges to a sequence of fixed points; after reading the para-RNN paper, how to traverse this set of fixed points efficiently in parallel became clear, forming FBT's technical path.
- He cited the Mozer et al. arXiv paper "Recirculation" as supporting evidence, which reported a 23% reduction in perplexity.
- Langford noted a parallel line of development: FBT is essentially recurrent training across the entire network depth (full-stack depth recurrence).
Unconfirmed
- "OpenAI's Astra project uses the Full Bandwidth Transformer" remains just a rumor. Langford explicitly stated he has no inside information and cannot confirm it; he added that if true, it would be parallel independent development.
- The only suspicious signal was someone from OpenAI asking where the idea came from, which Langford found strange — but this is not evidence that Astra uses FBT.
Why it matters
- If FBT can train smoothly without nextlat and leverage para-RNN's approach for efficient parallel traversal of fixed points, it could open a new path toward recurrent-style training of Transformer architectures; the 23% perplexity reduction in the Recirculation paper is a preliminary positive signal.
- OpenAI proactively asking about the idea's origin gives the rumor some credibility despite remaining unverified, sparking community discussion about how research paths at big labs and independent researchers converge.
2026-09-04 ~ 2026-09-04 · 5 related posts
Primary sources
- Rumor: OpenAI's Astra uses Full Bandwidth Transformer; Langford claims parallel development — JohnCLangford ·
- Langford explains FBT theory: nextlat's fixed-point convergence and para-RNN parallel training — JohnCLangford ·
- Langford on FBT surprises: no nextlat needed as Recirculation paper cuts perplexity 23% — JohnCLangford ·
- Langford: OpenAI staffer probed where the FBT idea came from amid Astra rumors — JohnCLangford · 2026-09-04
- [source] Rumor: OpenAI's Astra uses Full Bandwidth Transformer; Langford claims parallel development — JohnCLangford · 2026-09-04
- [source] Langford explains FBT theory: nextlat's fixed-point convergence and para-RNN parallel training — JohnCLangford · 2026-09-04
- Langford: Full Bandwidth Transformer drops nextlat and trains recirculation over full stack depth — JohnCLangford · 2026-09-04
- [source] Langford on FBT surprises: no nextlat needed as Recirculation paper cuts perplexity 23% — JohnCLangford · 2026-09-04