Prepending ".\n\n Okay" lifts Olmo-3-7B's MATH-500 accuracy from 42% to 78%, hinting base models already reason
arankomatsuzaki · x · 2026-10-06
arankomatsuzaki shares a striking finding: base models can already reason if given a cue drawn from their training data.
- Prepending ".\n\n Okay" to the prompt raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%.
- Similarly, " Alright ," lifts Qwen3-14B from 72% to 87%.
- The mechanism: RL makes such stylistic cues more likely; fixing them into the base model recovers much of RL's performance gain over the base.
The implication is that part of RL post-training's benefit may simply be eliciting reasoning behaviors already latent in the base model, rather than creating new capabilities.
Related event: Study: Two Opening Tokens Unlock Base Model Reasoning Without RL(4 posts)→
More from Models
- Are models only improving at verifiable domains? AI stories winning prizes spark debate — erikphoel · 2026-10-06
- Mistral fights back: 38 points at $1.13/task, best open model in the West, top cyber score — rickasaurus · 2026-10-06
- 8 models play Tetris side by side: Perplexity Decider tops decision benchmark — AravSrinivas · 2026-10-06
- Mistral's new 1T-param model 'Le Chonk' is a deliberate nod to an X meme — shaunralston · 2026-10-06
- Dev Says Opus 5.5 Is First Model He Trusts With Reasoning Turned Down From Max to High — casper_hansen_ · 2026-10-06
- ChatGPT voice mode keeps leaking its hidden system instructions to users — VoidStateKate · 2026-10-06