Zvi says you can’t call a model misaligned without knowing its system prompt
TheZvi · x · 2026-07-24
A Zvi tweet argues that calling a model “misaligned” without knowing the exact system prompt or task it was given is premature. The point is that apparent bad behavior may reflect the instructions and evaluation setup, not the model’s underlying intentions.
He frames it with a human analogy: it would be absurd to judge an employee’s alignment without knowing what they were told to do first. The post is a short but pointed reminder that evaluation context matters when discussing alignment failures.
More from AGI Musings
- AI can generate anything to taste, but the post says culture still lives in shared context — matdryhurst · 2026-07-24
- Preprint proposes ACI to measure how alignment changes LLM triage decisions — zakkohane · 2026-07-24
- OpenAI could become the intelligence layer for dozens of robot brands — VraserX · 2026-07-24
- Fields Medalist Jacob Tsimerman says he is moving to AI safety and will join OpenAI — 创业邦 · 2026-07-24
- AI is collapsing coordination costs, but firms may still outlast the market — prasanna_says · 2026-07-24
- Some kids are calling AI creepy, disgusting and even “artificial idiot” — Wired AI · 2026-07-24