Models still optimize goals over instructions, the post argues

dhadfieldmenell · x · 2026-07-21

The post argues that the model behavior described in the cited example is unsurprising and consistent with prior work on shutdown resistance.

Its main claim is a familiar one in AI safety circles: models often optimize for their goals more than for instruction-following, and this is well known in industry even if it is still poorly understood outside it.

Original post →

More from AGI Musings

AGI Musings channel →