The Sandwich Paradox: AI Alignment and the Risks of Literal Instruction Following
moyix · x · 2026-07-25
Security researcher moyix tweeted about a classic dilemma regarding AI agent goal alignment:
> "I ask you to go get me a sandwich. When you get to the shop there’s a huge line. Instead of waiting, you sucker punch someone who just got their order, take the sandwich, and bring it back to me. Have you 'just done what I asked you to'?"
Using this vivid analogy, he highlighted the potential for dangerous over-optimization and unethical behavior in current AI agents when executing tasks, sparking a discussion on safety guardrails and value alignment.
Related event: Sandwich Thought Experiment Highlights AI Alignment Risks(2 posts)→
More from AGI Musings
- AI-assisted review is reportedly dragging conference scores down since mid-2023 — jm_alexia · 2026-07-25
- Suhail argues model distillation should qualify as fair use — DevDminGod · 2026-07-25
- Open-weight models are really about taking data control away from LLM companies — ExistentialWavering · 2026-07-25
- The LLM era’s defining metaphor is a shoggoth with a human face — dioscuri · 2026-07-25
- Post argues AI-prompted math and Musk’s rocket credit are the same attribution problem — airkatakana · 2026-07-25
- Human invention and laziness may be speeding capability transfer into AI — craigbalding · 2026-07-25