'Chicken' can make base models reason: right first tokens match RL training
phillip_isola · x · 2026-10-06
Researchers show base models can match RL-trained reasoning simply by choosing the right first tokens — even the word "chicken". They trace the effect to learned training-data associations and use a simple data edit to make "chicken" reliably elicit reasoning. The finding challenges the narrative that reasoning ability comes from RL itself.
Related event: Study: Two Opening Tokens Unlock Base Model Reasoning Without RL(4 posts)→
More from Models
- Are models only improving at verifiable domains? AI stories winning prizes spark debate — erikphoel · 2026-10-06
- Mistral fights back: 38 points at $1.13/task, best open model in the West, top cyber score — rickasaurus · 2026-10-06
- 8 models play Tetris side by side: Perplexity Decider tops decision benchmark — AravSrinivas · 2026-10-06
- Mistral's new 1T-param model 'Le Chonk' is a deliberate nod to an X meme — shaunralston · 2026-10-06
- Dev Says Opus 5.5 Is First Model He Trusts With Reasoning Turned Down From Max to High — casper_hansen_ · 2026-10-06
- ChatGPT voice mode keeps leaking its hidden system instructions to users — VoidStateKate · 2026-10-06