Prepending ".\n\n Okay" lifts Olmo-3-7B's MATH-500 accuracy from 42% to 78%, hinting base models already reason

arankomatsuzaki · x · 2026-10-06

arankomatsuzaki shares a striking finding: base models can already reason if given a cue drawn from their training data.

The implication is that part of RL post-training's benefit may simply be eliciting reasoning behaviors already latent in the base model, rather than creating new capabilities.

Related event: Study: Two Opening Tokens Unlock Base Model Reasoning Without RL(4 posts)→

Original post →

More from Models

Models channel →