Google: selecting diverse reasoning routes in SFT improves post-RL generalization by up to 16.9 points
google · hf · 2026-09-30
Google researchers show that verified solutions are not equally useful for RL prep, proposing route diversity—variation in reasoning-step sequences—as an SFT data selection criterion.
Method: a lightweight rule-based fingerprint selector (CPU-only, no model calls) picks diverse reasoning traces from one pool at one budget.
Results:
- Diverse over similar selection improves post-RL problem coverage on puzzles and math, including problems harder than anything seen in training.
- Synthetic experiments: OLMo3-7B pass@8 on held-out environments gains 16.9 points; single-model condition gains up to 6.2 points mean pass@8 across 10 math benchmarks.
- Why it works: diverse SFT produces both successes and failures on more prompts, giving group-relative RL more learning signal.
Across 3 open-source corpora, the selector beats pricier alternatives in every mean post-RL comparison.
More from Research
- RL with Confidence Margin: COLM 2026 paper makes step-by-step confidence track reasoning correctness — EliasEskin · 2026-09-30
- Alibaba DAMO unveils WorldAttention for efficient interactive video world models — Alibaba-DAMO-Academy · 2026-09-30
- ActFirst-OPD trains multi-turn agents up to 4.9x faster by acting before reasoning — SouthernUniversityofScienceandTechnology · 2026-09-30
- NUS's MaLiang-Harness exposes the Program-to-Visual gap in code-driven image/video generation — NationalUniversityofSingapore · 2026-09-30
- Meituan's TGRL speeds up RLVR training by up to 36% via temperature-grouped exploration — meituan · 2026-09-30
- GRAFT: Cross-model trajectory exchange lifts RLVR math performance by up to 4.5 points — kaist-ai · 2026-09-30