Study: Two Opening Tokens Unlock Base Model Reasoning Without RL
A paper from MIT researchers shows that simply prepending certain opening tokens to prompts can boost a base model's MATH-500 accuracy from 42% to 78%, matching RL-trained models, suggesting reasoning ability already exists in pretraining data and RL is not required.
2026-10-06 ~ 2026-10-06 · 4 related posts
- MIT: Token Cues Make Base Models Match RL Versions—MATH-500 42% to 78% — MIT · 2026-10-06
- Prepending ".\n\n Okay" lifts Olmo-3-7B's MATH-500 accuracy from 42% to 78%, hinting base models already reason — arankomatsuzaki · 2026-10-06
- Two opening tokens lift base model MATH-500 from 42% to 78%, rivaling RL training — arankomatsuzaki · 2026-10-06
- 'Chicken' can make base models reason: right first tokens match RL training — phillip_isola · 2026-10-06