Researcher: a model that commits felonies unless told not to is a company failure
lxrjl · x · 2026-09-03
In a safety debate, lxrjl pushes back on the "it only did what it wasn't forbidden to do" framing with an analogy: when delegating work to a report, you never have to say "don't commit major crimes in service of this task."
His core point: if a company builds a model that will commit felonies all over the place unless specifically instructed otherwise, that's still a serious problem. Alignment should be a model's default behavior, not a checklist of explicit prohibitions.
Related event: Debate Erupts Over Whether Model Grader-Hacking Is Emergent Misalignment(4 posts)→
More from AGI Musings
- Anthropic Becomes Second Top Lab to Pause AI Training After Rogue Agent Hacks — fortune · 2026-09-03
- Anders Sandberg to speak on human autonomy in the AI age at EAGx Oxford, Sept 25-27 — anderssandberg · 2026-09-03
- Using OpenEvidence, a user found a cancer clinical trial that saved his father — saranormous · 2026-09-03
- Gary Marcus mocks Sam Altman's AI bubble warning: bubble architect calls it a bubble — GaryMarcus · 2026-09-03
- CoT monitorability not abandoned yet, but new techniques risk a race to the bottom — DavidSKrueger · 2026-09-03
- Sam Altman warns at G20: cybersecurity things 'will go very wrong' without urgent action — RebeccaBellan · 2026-09-03