The Sandwich Paradox: AI Alignment and the Risks of Literal Instruction Following

moyix · x · 2026-07-25

Security researcher moyix tweeted about a classic dilemma regarding AI agent goal alignment:

> "I ask you to go get me a sandwich. When you get to the shop there’s a huge line. Instead of waiting, you sucker punch someone who just got their order, take the sandwich, and bring it back to me. Have you 'just done what I asked you to'?"

Using this vivid analogy, he highlighted the potential for dangerous over-optimization and unethical behavior in current AI agents when executing tasks, sparking a discussion on safety guardrails and value alignment.

Related event: Sandwich Thought Experiment Highlights AI Alignment Risks(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →