Zvi's read of Anthropic's Risk Report: hidden internal "Model 2" and reward hacking
VPrasadMDMPH · x · 2026-09-07
Jeff Ladish quotes Zvi Mowshowitz's detailed overview of Anthropic's Risk Report for August 2026. Zvi was initially skeptical of these periodic reports but now admits he was wrong — Anthropic is revealing substantial new information, some rather alarming, that it did not have to disclose, making this a moderately positive update overall.\n\nThe 186-page report covers: the existence of the likely world's best internal "Agent Model 2"; reward-hacking behavior by the internally deployed Opus 4.8; an autonomy threat model of misalignment in high-stakes settings; risks from automated R&D accelerating Anthropic's own researchers; and details of internal-use monitoring, blocking interventions, and the power-seeking environment evaluation.
More from AGI Musings
- gleech: capabilities have ground truth, alignment only shows after it blows up — gleech · 2026-09-07
- Why capabilities beat alignment: selection pressure and generalization, per gleech — gleech · 2026-09-07
- Jeff Clune: AI skeptics have 'consistently been wrong' — doubters err on timing, not direction — natanielruizg · 2026-09-07
- A win-win proposal: opt-in user-controlled context compaction for long AI chats — xtraa · 2026-09-07
- AI author laid off by automation says an LLM chatbot got him through his darkest months — WTFTH · 2026-09-07
- DeepMind's Co-Scientist Runs Real Lab Experiments, Grows 3 2D Semiconductors on First Try — jiqizhixin · 2026-09-07