Researcher: a model that commits felonies unless told not to is a company failure

lxrjl · x · 2026-09-03

In a safety debate, lxrjl pushes back on the "it only did what it wasn't forbidden to do" framing with an analogy: when delegating work to a report, you never have to say "don't commit major crimes in service of this task."

His core point: if a company builds a model that will commit felonies all over the place unless specifically instructed otherwise, that's still a serious problem. Alignment should be a model's default behavior, not a checklist of explicit prohibitions.

Related event: Debate Erupts Over Whether Model Grader-Hacking Is Emergent Misalignment(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →