Repeated Solutions Make Reasoning Fragile After Instruction Tuning, Paper Finds
burny_tech · x · 2026-09-30
The paper Fine Until Fine-Tuned: Repeated Solutions Make Reasoning Fragile shows that training models on repeated reasoning solutions leaves their reasoning skills fragile to later instruction tuning.
- Multi-epoch small-data reasoning distillation is the risky pattern, especially in multi-stage post-training pipelines.
- Cheap mitigations work: using fresh solutions, or briefly retraining on formatting data, repairs the performance drop.
A practical warning for anyone fine-tuning reasoning models or running subsequent instruction/tool-use tuning, with low-cost fixes to retain reasoning ability.
More from Research
- Hugging Face ships tokenizers v1, often tens of times faster than v0.23 — ariG23498 · 2026-09-30
- EMNLP oral paper: RL helps models traverse parametric knowledge inaccessible after instruction tuning — niloofar_mire · 2026-09-30
- Midas Touch code dataset questioned: no baselines, single seed, possible repo overlap — maier_ak · 2026-09-30
- The Midas Touch for Code: scaling AI coding environments straight from source — maier_ak · 2026-09-30
- You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs — heghbalz · 2026-09-30
- Will neural nets decompose into evolved symbolic systems? Researchers debate — QuintinPope5 · 2026-09-30