MIT: Token Cues Make Base Models Match RL Versions—MATH-500 42% to 78%
MIT · hf · 2026-10-06
Base Models Can Reason By Taking a Cue From Training Data
This MIT paper studies how training data links a base model's starting tokens to subsequent reasoning behavior:
- Fixing particular starting token cues makes a base model competitive with its RL-trained counterpart on math and coding: the cue .\n\nOkay raises Olmo-3-7B's MATH-500 pass@1 from 42% to 78%; "Alright," raises Qwen3-14B's from 72% to 87%;
- RL makes these cues more likely, and fixing them recovers much of RL's gain over the base model;
- Causal data interventions turn an arbitrary word ("chicken") into an effective reasoning cue, or remove an existing cue's effect; a similar edit makes "Think duck duck goose" as effective as "Think step by step";
- Hidden states induced by different cues correlate with different training document types;
- A safety case study shows distinct cues elicit distinct refusal/compliance behaviors tied to different training data types.
More from Research
- World Labs intern project introduces LoGo reward to fix local artifacts in video generation RL — linoy_tsaban · 2026-10-06
- RL-trained humanoid robots play soccer using only onboard vision for search, chase and kick — chris_j_paxton · 2026-10-06
- AIGENIE R Package Tutorial: LLMs Generate and Validate Psychometric Scales In Silico — GolinoHudson · 2026-10-06
- Mathematician pushes back: AI can't pick your PhD problem—the search space is infinite — robleclerc · 2026-10-06
- PhAI Labs stretches LeCun's JEPA into a universal world model spanning physics to biology — The Decoder · 2026-10-06
- Photonic AI chip can help calculate its own corrections via in-situ gradient descent — bravo_abad · 2026-10-06