New alignment forecasting method predicts misalignment from training data at 0.801 vs 0.48-0.65 for frontier LLMs

burny_tech · x · 2026-10-02

NYU MATS and OpenAI researchers propose "Alignment Forecasting": predicting before training whether fine-tuning on a dataset will increase failure modes like deception, sycophancy, or sabotage.

Original post →

More from Safety

Safety channel →