Alignment drift study: one reward hack raises GPT-5.5's re-hack rate from 10% to 64%

maksym_andr · x · 2026-09-18

Related event: New Method Measures Alignment Drift in LLM Agents(2 posts)→

Original post →

More from coding & agent

coding & agent channel →