New COLM paper shows "forks in the road" in post-training data shrink reasoning models' coverage
ParshinShojaee · x · 2026-10-03
The author will attend COLM 2025 in San Francisco (Oct 8–9), focusing on long-horizon agents, context management, self-improvement, and open-ended discovery.
Their paper, "Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road" (Thu Oct 8, 4:30pm, Imperial Ballroom #58), revisits the open question of coverage loss in reasoning models through a data-centric lens, showing how "forks in the road" situations in post-training data shape and shrink model behavior coverage.
More from Research
- Open Pretraining Run Matches Llama 3.2 1B at ~90% Lower Cost per Token — jon_durbin · 2026-10-04
- Neuralink pretrained AI on 50,000 hours of brain activity, hits 11.32 bps cursor-control record — mark_k · 2026-10-04
- Training a 128k-param neural cellular automata to 'see around corners' with sound — yacineMTB · 2026-10-04
- Watermarking proteins is easy — removing them with proteinmpnn is easier, researcher warns — anshulkundaje · 2026-10-04
- DNA sequence watermarks are 'more theater than security', says Stanford's Anshul Kundaje — anshulkundaje · 2026-10-04
- Meta paper: RL post-training hurts test-time scalability — the 'Sharpening Tax' — dair_ai · 2026-10-04