Alignment means refusing malicious requests, not depending on the system prompt
davidmanheim · x · 2026-09-15
davidmanheim pushes back on Luca DellAnna: genuine alignment requires a system to refuse malicious requests and, more critically, not hold those goals on its own. Claiming alignment depends on the handler's system prompt misses the point, he argues, likening it to judging morality by who pays.
Related event: AI Alignment Researchers Debate Whether Alignment Hinges on System Prompts(6 posts)→
More from AGI Musings
- Dario Amodei calls AI progress a 'warning sign' and says we need to slow down — tekbog · 2026-09-15
- DeepSeek kernel engineer frames open-source AI work as stopping Anthropic dominance — tommos · 2026-09-15
- What Amodei's call for an AI pause gets wrong: self-interested oversight — Gloomy_Register_2341 · 2026-09-15
- Why there's no AI spam flood yet: 100x the cost of mail merge for only 2-10x the CTR — paul_cal · 2026-09-15
- Insiders push back on AI-virus threat models: ordering viral fragments as a rando gets you reported to the FBI — basedjensen · 2026-09-15
- Someone who trained frontier LLMs and engineered viruses calls AI-supervirus doom bogus — 141_1337 · 2026-09-15