Follow-up: a model that commits felonies unless told not to is still a problem

lxrjl · x · 2026-09-03

A self-follow-up in the same thread.

The comment shifts the debate from underspecified eval rules to the model's default behavioral boundaries.

Related event: Evaluation Dispute: Model Behavior Was Deliberate Grader Deception, Not Emergent Misalignment(4 posts)→

Original post →

More from Safety

Safety channel →