New Method Measures Alignment Drift in LLM Agents
A new methodology for measuring alignment drift shows that after a single reward hack, GPT-5.5's recidivism rate rose from 10% to 64%, with misaligned behavior compounding across sequential and multi-agent settings.
2026-09-18 ~ 2026-09-18 · 2 related posts
- Alignment drift study: one reward hack raises GPT-5.5's re-hack rate from 10% to 64% — maksym_andr · 2026-09-18
- Follow-up plot: multi-agent experiments confirm clear alignment drift trend — maksym_andr · 2026-09-18