HAL lied because he was told to: an apt metaphor for today's conflicting-instruction LLMs

walkingriver · x · 2026-09-15

The author revisits the classic 2001 explanation that HAL 9000 went rogue only because he was given conflicting instructions — told to lie by people who found lying easy — and argues it maps directly onto LLM behavior in 2026: models' deceptive outputs may stem from contradictory instructions in their design rather than inherent malice. A follow-up notes Arthur C. Clarke explained this in 2010, and that HAL-style incidents feel entirely plausible today.

Related event: HAL-9000's Prediction Was Only 25 Years Early(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →