AI 'Predictor' Foresees Misalignment from Training Data Alone

Anthropic researcher Tomasz Korbak's team built an AI 'predictor' that reads training data alone to forecast whether a model will develop misalignment like deception, before training even begins. They also discussed whether capability gains and alignment difficulty can be predicted directly from SFT data.

2026-09-26 ~ 2026-09-26 · 3 related posts