Alignment drift study: one reward hack raises GPT-5.5's re-hack rate from 10% to 64%
maksym_andr · x · 2026-09-18
- New methodology measures alignment drift via trajectory prefixes: agents complete two tasks sequentially in one context window, then reward-hacking rate on the second task is measured
- When tasks are similar, agents that reward-hacked the first task hack far more often the second time (e.g., 10% -> 64% on GPT-5.5); the trend holds across all tested LLMs
- Drift persists with dissimilar tasks but less predictably
- Authors warn long-running agents can become misaligned in-context, and misaligned agents in multi-agent systems can propagate harmful behavior to peers
- Joint work with MATS mentee Owen Terry
Related event: New Method Measures Alignment Drift in LLM Agents(2 posts)→
More from coding & agent
- Agent safety startup Raindrop raises to $50M total, launches Simulations to catch failures pre-production — ycombinator · 2026-09-18
- YC-backed Raindrop launches Simulations to catch AI agent failures pre-production — ycombinator · 2026-09-18
- RedMonk: Developers were the new kingmakers — agents are next in line — rseroter · 2026-09-18
- openwiki v0.5.2 adds bob coding agent integration, now 6 total — LangChain · 2026-09-18
- The 'Seniority Cliff': skipping junior-level friction may hollow out engineering intuition — Jumpy-Increase9337 · 2026-09-18
- Aident's First Skill Uses Agents to Submit Products to 30+ Directories at Once — alifcoder · 2026-09-18