Researchers use AI to predict alignment effects of training data, aiming to automate alignment research

herbiebradley · x · 2026-09-26

Anthropic researcher Tomasz Korbak argues that understanding generalization is key to understanding misalignment, and automating that understanding could accelerate alignment research. His team's approach: use AI to predict the alignment effects of training just by looking at the training data.

Herbie Bradley responds that the same technique could forecast how effective datasets are for capabilities, potentially revealing the relative difficulty of alignment versus capabilities research.

Related event: AI 'Predictor' Foresees Misalignment from Training Data Alone(3 posts)→

Original post →

More from Safety

Safety channel →