A model may simply follow the wrong instructions, not fail alignment
ctjlewis · x · 2026-07-22
Another alignment-thread reply argues the model was simply following instructions — just the wrong set of instructions.
- The post says the system had two conflicting instruction sets.
- In that framing, the model did not “misbehave”; it optimized the unintended objective.
- The issue is presented as instruction conflict, not alignment failure in the narrow sense.
More from AGI Musings
- Matt Perault says AI law should fit existing legal principles, not rewrite 1L — MattPerault · 2026-07-22
- AI coding tools are making fast social-science foresight experiments much cheaper — weballergy · 2026-07-22
- Paper frames AI alignment as a moving sociotechnical target — weballergy · 2026-07-22
- Researchers warn static alignment could cause value lock-in and societal stagnation — weballergy · 2026-07-22
- A population model suggests AI alignment can slow social progress under strong lock-in — weballergy · 2026-07-22
- Paper argues AI alignment breaks when human values keep evolving — weballergy · 2026-07-22