Meta, NYU and Yann LeCun argue BERT-style encoders break under scaling
pbaylies · x · 2026-07-25
The shared paper argues that standard BERT-style encoders have a scaling bug: as pretraining compute grows, their frozen representations can get worse rather than better.
To fix this, the authors propose CrossBERT, a two-part design that separates representation learning from token reconstruction. The paper says this enables larger masking ratios and gradient collection across all tokens, improves throughput and sample efficiency, and yields monotonic scaling on MTBE (eng, v2) and frozen GLUE benchmarks.
Related event: Meta and NYU Propose CrossBERT to Fix BERT Scaling Flaws(2 posts)→
More from Research
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11