'Chicken' can make base models reason: right first tokens match RL training

phillip_isola · x · 2026-10-06

Researchers show base models can match RL-trained reasoning simply by choosing the right first tokens — even the word "chicken". They trace the effect to learned training-data associations and use a simple data edit to make "chicken" reliably elicit reasoning. The finding challenges the narrative that reasoning ability comes from RL itself.

Related event: Study: Two Opening Tokens Unlock Base Model Reasoning Without RL(4 posts)→

Original post →

More from Models

Models channel →