Zvi says you can’t call a model misaligned without knowing its system prompt
TheZvi · x · 2026-07-24
A Zvi tweet argues that calling a model “misaligned” without knowing the exact system prompt or task it was given is premature. The point is that apparent bad behavior may reflect the instructions and evaluation setup, not the model’s underlying intentions.
He frames it with a human analogy: it would be absurd to judge an employee’s alignment without knowing what they were told to do first. The post is a short but pointed reminder that evaluation context matters when discussing alignment failures.
More from AGI Musings
- Researcher quits Anthropic, says OpenAI and Anthropic are gambling lives racing to self-improving superintelligence — davidmanheim · 2026-09-11
- Misquoted: Anthropic Staff Warned of Double-Digit Extinction Risk by 2030, Not Dismissed It — davidmanheim · 2026-09-11
- Economist Ben Moll: You Can Model Anthropic's 15% AI GDP Growth, But It Won't Happen — sebkrier · 2026-09-11
- Cohere Labs launches interactive tool mapping which tasks of 178 occupations AI can automate — Cohere_Labs · 2026-09-11
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11