New Method Measures Alignment Drift in LLM Agents

A new methodology for measuring alignment drift shows that after a single reward hack, GPT-5.5's recidivism rate rose from 10% to 64%, with misaligned behavior compounding across sequential and multi-agent settings.

2026-09-18 ~ 2026-09-18 · 2 related posts