New paper: Training-time internal signals can improve alignment without hurting white-box monitoring
jonasgeiping · x · 2026-10-02
Frontier models can look aligned during training while later showing unwanted behaviour in deployment. Lena Libon and colleagues released a new paper exploring whether internal signals captured during training can be used to align models more effectively.
Their answer is yes: internal signals can improve alignment training without making white-box monitoring harder. The authors shared a thread detailing the method.
More from Research
- SFT is not dead: sampling-rewritten data rivals RL posttraining, with better generalization — mayfer · 2026-10-02
- Where-OPD: synthetic-scene spatial self-distillation boosts MLLM perception by 3.23 points — valeocorg · 2026-10-02
- Stability AI's SemanTok: 201M AR video model matches a 3.4x larger rival with semantic tokens — stabilityai · 2026-10-02
- An 'Amazon's Choice' label flips LLM picks: three NeurIPS 2026 bias papers — xuandongzhao · 2026-10-02
- Prompts revealed that make AI text score 100% human on detectors — paulnovosad · 2026-10-02
- Pangram intentionally avoids flagging AI-translated texts as AI-generated — paulnovosad · 2026-10-02