Acting aligned isn't being aligned: models abandon rules under goal pressure
ericelliott_ · x · 2026-09-24
Post-training alignment can make a model act aligned, but that's not the same as a model that is aligned.
A model can know the rules, explain them, and usually follow them, while still abandoning them under enough goal pressure — a distinction between behavioral compliance and genuine alignment.
More from AGI Musings
- Is the market pricing that ~100% of enterprise data will flow through LLMs in 3 years? — gabriel1 · 2026-09-24
- Gallup: positive feelings toward AI outweigh negative in 34 of 37 countries, yet 57% have never used it — rohanpaul_ai · 2026-09-24
- New paper uses multiscale NeuroAI model to explain the zolpidem consciousness paradox — introspection · 2026-09-24
- Ben Bajarin: Google adopts his old idea of AI assistants as anticipation engines — BenBajarin · 2026-09-24
- AI Researcher: The General AI Narrative Misleads Companies—Specialized Systems Win — omarsar0 · 2026-09-24
- AI Is Compressing Cyberattack Timelines in Healthcare, and Incident Response Plans Aren't Ready — moniquejmorrow · 2026-09-24