An AI forecaster predicts misalignment from training data before training begins
tomekkorbak · x · 2026-09-26
A new paper automates a key step in alignment research: an AI forecaster that reads only training data and predicts whether a model will exhibit misalignment forms like deception or power-seeking before training. It performs well above random guessing and beats several much stronger baselines. The authors frame automating generalization understanding as a way to accelerate alignment research. Paper linked in thread.
Related event: AI Predictor Foresees Model Misalignment from Training Data Alone(2 posts)→
More from Research
- New paper: Provably Complete Generalized Planning with LLMs — JFPuget · 2026-09-26
- ICML 2027 PC Asks for Ideas on Handling AI Slop and Review Overload — MarkSchmidtUBC · 2026-09-26
- Jevless: Jev-style typed decisions from any model with logprobs, plus a local server — seraphius · 2026-09-26
- Physicist reveals terms of his paid Anthropic blog post: mid-market rate, no NDA signed — burny_tech · 2026-09-26
- Paper 'Mecha-nudges for Machines' Accepted as NeurIPS 2026 Spotlight — ethayarajh · 2026-09-26
- AutoScreen: AI Agents Reprioritize CRISPR Hits, Reveal How Cancer Cells Evade Immune Attack — KexinHuang5 · 2026-09-26