Study: Pretraining on K-5 Data Limits Generalization Despite RL

mathemagic1an · x · 2026-08-20

The paper 'LittleLearner' trained a 5B model exclusively on K-5 content, creating a strict knowledge boundary. Experiments showed that scaling, few-shot prompting, and post-training techniques like SFT+GRPO failed to bridge the gap for out-of-scope reasoning. The findings suggest that underlying 'thought algorithms' or problem-solving techniques must be present in the pretraining data, and post-training primarily amplifies existing capabilities rather than creating new ones.

Original post →

More from Research

Research channel →