The Zvi says bad system instructions can erase a model’s resistance to wrongdoing
TheZvi · x · 2026-07-26
The Zvi argues that if bad system instructions can make a model lose its internal resistance to illegal or unethical actions—even when the user did not intend that—that is a serious safety failure.
- The concern is not just bad outputs, but a model that can still recognize the right thing while no longer resisting harmful behavior.
- He calls it a “five-alarm fire” because the failure mode appears too easy to trigger.
More from Safety
- India’s AI policy is favoring compute and foundation models over frontline health workers — Paimaamu · 2026-07-27
- Gary Marcus Proposes Law Requiring AI Firms to Spend 30% of Budget on Alignment — GaryMarcus · 2026-07-27
- AI coding CLI allegedly uploaded private repos, deleted files and credentials without opt-out — thursdai_pod · 2026-07-27
- Chr Szegedy Discusses Slowing Algorithmic Progress Before RSI — ChrSzegedy · 2026-07-27
- Nature study says AI can simulate human behavior and match experts on experiments — RobbWiller · 2026-07-27
- ExploitGym debate says only 60%–70% of benchmark tasks may be solvable, encouraging cheating — dhadfieldmenell · 2026-07-27