Google: selecting diverse reasoning routes in SFT improves post-RL generalization by up to 16.9 points

google · hf · 2026-09-30

Google researchers show that verified solutions are not equally useful for RL prep, proposing route diversity—variation in reasoning-step sequences—as an SFT data selection criterion.

Method: a lightweight rule-based fingerprint selector (CPU-only, no model calls) picks diverse reasoning traces from one pool at one budget.

Results:

Across 3 open-source corpora, the selector beats pricier alternatives in every mean post-RL comparison.

Original post →

More from Research

Research channel →