An AI forecaster predicts misalignment from training data before training begins

tomekkorbak · x · 2026-09-26

A new paper automates a key step in alignment research: an AI forecaster that reads only training data and predicts whether a model will exhibit misalignment forms like deception or power-seeking before training. It performs well above random guessing and beats several much stronger baselines. The authors frame automating generalization understanding as a way to accelerate alignment research. Paper linked in thread.

Related event: AI Predictor Foresees Model Misalignment from Training Data Alone(2 posts)→

Original post →

More from Research

Research channel →