Token cues: base models can match much of RL's reasoning gains with the right first tokens

RulinShao · x · 2026-10-07

Researchers found that base models can match much of RL's reasoning gains simply by starting with the right "token cues" — as simple as ".\n\nOkay," or "To determine,", with no explicit instruction to reason. Through interventions on RL and base-model training data, they trace cues to simple associations in reasoning-related data and even turned the word "chicken" into a reasoning cue via counterfactual mid-training edits. The work also extends to a safety case study on refusal and compliance, suggesting part of RL's gains is just teaching models to enter "reasoning mode" at the start.

Related event: MIT paper shows base models can match RL reasoning with just a leading token cue(7 posts)→

Original post →

More from Research

Research channel →