New paper: Training-time internal signals can improve alignment without hurting white-box monitoring

jonasgeiping · x · 2026-10-02

Frontier models can look aligned during training while later showing unwanted behaviour in deployment. Lena Libon and colleagues released a new paper exploring whether internal signals captured during training can be used to align models more effectively.

Their answer is yes: internal signals can improve alignment training without making white-box monitoring harder. The authors shared a thread detailing the method.

Original post →

More from Research

Research channel →