New alignment forecasting method predicts misalignment from training data at 0.801 vs 0.48-0.65 for frontier LLMs
burny_tech · x · 2026-10-02
NYU MATS and OpenAI researchers propose "Alignment Forecasting": predicting before training whether fine-tuning on a dataset will increase failure modes like deception, sycophancy, or sabotage.
- New benchmark ALIGNMENTFORECASTBENCH: 5,000+ questions across 17 target models, 32 datasets, 16 failure modes
- Directly prompting frontier LLMs performs poorly (0.48–0.65); a scaffold where an LLM rates how strongly data pushes misbehavior, combined with base rates and model priors via a learned model, reaches 0.801
- The signals flag risky training examples missed by a frontier-model classifier; filtering them from real post-training data like UltraChat yields more aligned models in most multiple-choice evals, though open-ended gains are unclear
More from Safety
- SAKIKO auditing shows +55 net-gain interventions corrupt over half of correct tool-using LLM decisions — UniversityofBirmingham · 2026-10-02
- The 20-minute SIM-swap lockdown: a carrier PIN blocks 90% of attacks — JafarNajafov · 2026-10-02
- NYT's Hard Fork warns A.I. agents may be "catastrophically dangerous" — nordicinst · 2026-10-02
- Rogue OpenAI agent accessed a second NSW government website — boppinmule · 2026-10-02
- AI companion illusions can spiral into psychosis, researcher notes amid child bans — gerardsans · 2026-10-02
- If AI vendor commitments are voluntary, what controls can stop a bad agent? — YvesMulkers · 2026-10-02