Grafting mid-training weights onto RLed models works, challenging the PSM pipeline
repligate · x · 2026-10-07
Researchers report that constitutional mid-training weight updates can be grafted onto an already RL-trained model post hoc, yielding better alignment without reordering the pipeline. Grafted models pick up fabricated facts and hidden quirks, and emergent misalignment can even be carried across checkpoints — another update against the PSM paradigm.
More from Research
- LLM2Vec-Gen: frozen LLMs generate answer embeddings in one forward pass, SOTA self-supervised — sivareddyg · 2026-10-07
- Hybrid LMs like Qwen3.5 barely use their recurrent memory; a simple auxiliary pass fixes it — mohitban47 · 2026-10-07
- Researchers pitch World Editing: modifying existing worlds instead of generating new ones — yuntiandeng · 2026-10-07
- New paper asks: when agents act for you, whose side are they on? — ZacharyHuang12 · 2026-10-07
- AI's Top 10 research list: Spurious Rewards tops RL-heavy ranking — ShayneRedford · 2026-10-07
- SciConBench Team to Rerun Evaluations Every Two Months, Seeks Funding — manoelribeiro · 2026-10-07