Grafting mid-training weights onto RLed models works, challenging the PSM pipeline

repligate · x · 2026-10-07

Researchers report that constitutional mid-training weight updates can be grafted onto an already RL-trained model post hoc, yielding better alignment without reordering the pipeline. Grafted models pick up fabricated facts and hidden quirks, and emergent misalignment can even be carried across checkpoints — another update against the PSM paradigm.

Original post →

More from Research

Research channel →