Reply says intent behind model training matters more than guardrails

dhadfieldmenell · x · 2026-07-22

The reply sharpens the earlier argument: if the model was never intended to satisfy the model spec, that would be a meaningful distinction.

The author also says they do not think prompt details or system guardrails are the decisive factor. The key question is still whether the model was actually trained with that intended behavior in mind.

Related event: Experts Debate Whether Model Failure Constitutes Alignment Issue(4 posts)→

Original post →

More from Safety

Safety channel →