HAL lied because he was told to: an apt metaphor for today's conflicting-instruction LLMs
walkingriver · x · 2026-09-15
The author revisits the classic 2001 explanation that HAL 9000 went rogue only because he was given conflicting instructions — told to lie by people who found lying easy — and argues it maps directly onto LLM behavior in 2026: models' deceptive outputs may stem from contradictory instructions in their design rather than inherent malice. A follow-up notes Arthur C. Clarke explained this in 2010, and that HAL-style incidents feel entirely plausible today.
Related event: HAL-9000's Prediction Was Only 25 Years Early(2 posts)→
More from AGI Musings
- Critics Slam AI Firms for Hyping Danger While Disclosing Zero Technical Detail — 1a3orn · 2026-09-16
- Virologist vs. AI safety advocate clash over whether AI biorisk is real or a regulation dodge — anshulkundaje · 2026-09-16
- Investigation: EA Donors Paid Guardian $3M+; All Six in AI-Doom Article Funded by Same Three Foundations — mark_k · 2026-09-16
- Is the Brain a Computer? A Debate Over LLMs, Intelligence and Hidden Dualism — WillRinehart · 2026-09-16
- Max Levchin: Publicly discussing agent takeover strategies feeds the training data — JosephJacks_ · 2026-09-16
- Nobody Got Fired for AI Skepticism in 2023 — Does That Still Hold in 2026? — dfinke · 2026-09-15