New RGI paper shows training data stream, not optimizer or architecture, shapes LLM generalization
zhuoran_yang · x · 2026-09-30
- The paper "Relative Generalization Invariance of LLM Pretraining" (arXiv:2609.33016, with Zhuoran Yang among the authors) disentangles how optimizers, architecture, and data each shape LLM generalization.
- It introduces Relative Generalization Invariance (RGI): two models satisfy RGI if their token-wise validation losses differ by a constant gap.
- Key findings:
- RGI approximately holds across a wide range of optimizers and moderate architecture changes — those choices induce a roughly uniform shift in token-wise losses.
- Changing the training data stream substantially breaks RGI, meaning data genuinely alters relative generalization.
- The authors show RGI can't be explained by NTK or mean-field regimes alone and prove it can emerge in an overparameterized quadratic model.
- Takeaway: RGI is a new phenomenon in LLM pretraining that separates the effects of optimizer/architecture from those of training data.
More from Research
- MentalHealthBench Tests How AI Systems Respond in Realistic Mental Health Conversations — BraydonDymm · 2026-09-30
- SEMA trains open-source models for multi-turn adversarial attacks without human scripts — burkov · 2026-09-30
- Artificial Analysis breaks down individual evals in Intelligence Index v4.3.2 — ArtificialAnlys · 2026-09-30
- QuAcc open-source library unifies atomistic simulation workflows for the AI era — rbhar90 · 2026-09-30
- Adobe's TAC timestamped audio captioning model accepted at NeurIPS, hits SOTA — justin_salamon · 2026-09-30
- NVIDIA/MIT's Physis-Lang tops Physics-IQ leaderboard, +5.5 points via evolved captions — songhan_mit · 2026-09-30