Study: Two Opening Tokens Unlock Base Model Reasoning Without RL

A paper from MIT researchers shows that simply prepending certain opening tokens to prompts can boost a base model's MATH-500 accuracy from 42% to 78%, matching RL-trained models, suggesting reasoning ability already exists in pretraining data and RL is not required.

2026-10-06 ~ 2026-10-06 · 4 related posts