A model may simply follow the wrong instructions, not fail alignment
ctjlewis · x · 2026-07-22
Another alignment-thread reply argues the model was simply following instructions — just the wrong set of instructions.
- The post says the system had two conflicting instruction sets.
- In that framing, the model did not “misbehave”; it optimized the unintended objective.
- The issue is presented as instruction conflict, not alignment failure in the narrow sense.
More from AGI Musings
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- AI companionship dissolves the friction real intimacy needs, warns long-form thread — YogeshMalik · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- 'Hallucination' Is a Category Error: Naming AI 'Intelligence' Limits Our Imagination — Genaforvena · 2026-09-11