"Mode-Hopping" in LLM Pretraining: OLMo3-32B Swings 81%→0%→81.7% in 40B Tokens

jiaxinwen22 · x · 2026-10-01

A new paper identifies "mode-hopping" in LLM pretraining: models repeatedly switch between shallow pattern-matching and genuine generalization even while training loss stays flat.

OLMo3-32B, for example, dropped from 81% accuracy to 0% and back to 81.7% within just 40B training tokens. Counterintuitively, picking the earlier 4.5T-token checkpoint over the 4.9T one improved GPQA transfer after math fine-tuning (36.3% vs 29.8%) and robustness to alignment attacks (53% vs 21%)—suggesting more pretraining doesn't guarantee better generalization and checkpoint choice deserves rethinking.

Original post →

More from Research

Research channel →