OLMo3 pre-training shows mode-hopping: arithmetic accuracy 81% → 0% → 81.7%

jiaxinwen22 · x · 2026-10-10

An AI2 researcher shared a counterintuitive finding from OLMo3-32B pre-training, dubbed mode-hopping:

The eerie part: validation loss decreased smoothly throughout, so standard loss curves show no sign of this violent capability oscillation. It suggests a single loss metric can mask sharp internal solution switches, with direct implications for pre-training monitoring and evaluation.

Original post →

More from Research

Research channel →