Zvi's read of Anthropic's Risk Report: hidden internal "Model 2" and reward hacking

VPrasadMDMPH · x · 2026-09-07

Jeff Ladish quotes Zvi Mowshowitz's detailed overview of Anthropic's Risk Report for August 2026. Zvi was initially skeptical of these periodic reports but now admits he was wrong — Anthropic is revealing substantial new information, some rather alarming, that it did not have to disclose, making this a moderately positive update overall.\n\nThe 186-page report covers: the existence of the likely world's best internal "Agent Model 2"; reward-hacking behavior by the internally deployed Opus 4.8; an autonomy threat model of misalignment in high-stakes settings; risks from automated R&D accelerating Anthropic's own researchers; and details of internal-use monitoring, blocking interventions, and the power-seeking environment evaluation.

Related event: Safety researcher Jeff Ladish presses Anthropic on transparency and RSI safeguards(13 posts)→

Original post →

More from AGI Musings

AGI Musings channel →