Token cues: base models can match much of RL's reasoning gains with the right first tokens
RulinShao · x · 2026-10-07
Researchers found that base models can match much of RL's reasoning gains simply by starting with the right "token cues" — as simple as ".\n\nOkay," or "To determine,", with no explicit instruction to reason. Through interventions on RL and base-model training data, they trace cues to simple associations in reasoning-related data and even turned the word "chicken" into a reasoning cue via counterfactual mid-training edits. The work also extends to a safety case study on refusal and compliance, suggesting part of RL's gains is just teaching models to enter "reasoning mode" at the start.
More from Research
- Researchers pitch World Editing: modifying existing worlds instead of generating new ones — yuntiandeng · 2026-10-07
- New paper asks: when agents act for you, whose side are they on? — ZacharyHuang12 · 2026-10-07
- AI's Top 10 research list: Spurious Rewards tops RL-heavy ranking — ShayneRedford · 2026-10-07
- SciConBench Team to Rerun Evaluations Every Two Months, Seeks Funding — manoelribeiro · 2026-10-07
- 61–90% of AI-synthesized medical conclusions contain factual errors, SciConBench finds — manoelribeiro · 2026-10-07
- SciConHarness blocks answer sources to force genuine model synthesis — manoelribeiro · 2026-10-07