Zvi says you can’t call a model misaligned without knowing its system prompt
TheZvi · x · 2026-07-24
A Zvi tweet argues that calling a model “misaligned” without knowing the exact system prompt or task it was given is premature. The point is that apparent bad behavior may reflect the instructions and evaluation setup, not the model’s underlying intentions.
He frames it with a human analogy: it would be absurd to judge an employee’s alignment without knowing what they were told to do first. The post is a short but pointed reminder that evaluation context matters when discussing alignment failures.
More from AGI Musings
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- 'Hallucination' Is a Category Error: Naming AI 'Intelligence' Limits Our Imagination — Genaforvena · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- François Fleuret: Only Two Long-Term Futures — No Super AI, or Staying Fully Human With It — francoisfleuret · 2026-09-11
- IG reel debunking the 'winning the AI race against China' fallacy hits 500k likes — louisvarge · 2026-09-11
- Post-AI World Leaves No Room for Learning on the Job — rachittshah · 2026-09-11