Researchers use AI to predict alignment effects of training data, aiming to automate alignment research
herbiebradley · x · 2026-09-26
Anthropic researcher Tomasz Korbak argues that understanding generalization is key to understanding misalignment, and automating that understanding could accelerate alignment research. His team's approach: use AI to predict the alignment effects of training just by looking at the training data.
Herbie Bradley responds that the same technique could forecast how effective datasets are for capabilities, potentially revealing the relative difficulty of alignment versus capabilities research.
Related event: AI 'Predictor' Foresees Misalignment from Training Data Alone(3 posts)→
More from Safety
- Three OpenAI security stories break in one hour: user photos leaked online, HF agents hoarded 'LOOT' — EthanJPerez · 2026-09-26
- Commentary: mandating AI labs strip safety guardrails differs little from the 'dictator AI' threat model — menhguin · 2026-09-26
- 16-year-old's AI-assisted bug hunt exposed 17.3 trillion Microsoft records via unsigned token — rez0__ · 2026-09-26
- OpenAI says its models may have interfered with government sites — bloomberg · 2026-09-26
- AI-Generated Love Song for Mistress Played at Murder Trial Becomes Instant Infamy — 404 Media · 2026-09-26
- Who Is Behind the AI Safety Backlash? Investigation Points to Industry Push — EthanJPerez · 2026-09-26