COLM paper: "forks in the road" in post-training data shrink reasoning model coverage
ParshinShojaee · x · 2026-10-09
Paper: Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road
- Reasoning models improve pass@1 via SFT post-training, but often show pass@k degradation vs the base model (coverage shrinkage).
- Hypothesis: shrinkage is driven by decision points or "forks in the road" in fine-tuning data — indecipherable patterns with multiple valid reasoning paths.
- Controlled case studies (indecipherable nodes in graph branching, reasoning modes) show shrinkage correlates tightly with the prevalence of such decision-point scenarios in training data.
- Mitigations: targeted data synthesis of decision points and a diversity-encouraging decoding mechanism.
- Conclusion: data-centric factors are a key driver of coverage shrinkage.
More from Models
- Wanted to bench Clef as an LLM judge, found a ton of broken rubrics instead — xeophon · 2026-10-09
- Blogger slams Claude's guardrails: female bodies must be veiled even in nonsexual contexts — flowersslop · 2026-10-09
- Artificial Analysis to launch Intelligence Index v5 in late October with new coding benchmark — ArtificialAnlys · 2026-10-09
- Nace launches NDI 1.0, a small doc-processing model claiming frontier-level parsing at 1/10th cost — nischay_twt · 2026-10-09
- Gemini 4 Argon appears in Google's own model picker ahead of keynote — vedantmisra · 2026-10-09
- ChatGPT Desktop Makes Itself Default CSV Reader, Users Call It Plainly Wrong — generativist · 2026-10-09