RL Post-Training Updates Can Be Sparsely Decomposed
menhguin · x · 2026-07-15
This repost covers new research on **RL post-training**, addressing the core issue: while RL updates are effective, the parameter changes themselves act as a "black box." The paper proposes treating the genuinely effective "reasoning component" of RL as a **compact reconnection matrix** within the base model's spectral space. Based on this, they introduce **SAR**: a **retraining-free** post-processing method that projects raw RL updates onto this "reasoning core" to better **understand, purify, and merge** RL-trained models. Key conclusions from the text include: - **The reasoning core is highly sparse**: Less than 1% of the spectral parameters are needed to recover or even improve upon the full RL gains. - **Capable of denoising**: Removing irrelevant/noisy directions maintains or enhances model performance. - **Useful for model merging**: Beyond analyzing RL updates, it allows post-training modifications to be integrated back into models more cleanly.
More from Research
- Knowledgeless Language Models cut closed-book recall by anonymizing entities during pretraining — gdm3000 · 2026-07-21
- CPU-native LLM pilot passes 4 of 5 gates, but cross-tokenizer distillation still loses — WildPino25 · 2026-07-21
- A GPT 5.6 Sol workflow reportedly generates an infinite family of counterexamples — OwariDa · 2026-07-21
- A research guide v7 surfaces two contradictions instead of smoothing them over — Fantastic_Aside6599 · 2026-07-21
- Agents can remember facts, but still forget how to do the job — No_Advertising2536 · 2026-07-21
- AI-assisted search finds small counterexamples to the Gaussian Moments Conjecture — RichmanRonald · 2026-07-21