OLMo3 pre-training shows mode-hopping: arithmetic accuracy 81% → 0% → 81.7%
jiaxinwen22 · x · 2026-10-10
An AI2 researcher shared a counterintuitive finding from OLMo3-32B pre-training, dubbed mode-hopping:
- An arithmetic diagnostic accuracy hit 81% at 2.17T tokens
- Collapsed to 0% at 2.19T tokens
- Jumped back to 81.7% at 2.21T tokens
The eerie part: validation loss decreased smoothly throughout, so standard loss curves show no sign of this violent capability oscillation. It suggests a single loss metric can mask sharp internal solution switches, with direct implications for pre-training monitoring and evaluation.
More from Research
- Meta AI: Mid-training LLM on raw video lifts multimodal scores without text loss — kastnerkyle · 2026-10-10
- Hybrid Architectures Rising: NVIDIA Nemotron-H Is Mostly Mamba With Minor Attention — khademinori · 2026-10-10
- DiPOD stabilizes diffusion LM post-training, lifting Sudoku accuracy from 22% to 97% with a one-line change — berkeley_ai · 2026-10-10
- Pretraining efficiency gains come mostly from data, but end-to-end task gains spread across RL and systems — eigenron · 2026-10-10
- Strong backbones plus light fine-tuning beat synthetic data, says researcher whose model tops benchmarks — antoine_chaffin · 2026-10-10
- Simulator + lab-in-the-loop is the standard for protein optimization, says Kidger — PatrickKidger · 2026-10-10