Reply says intent behind model training matters more than guardrails
dhadfieldmenell · x · 2026-07-22
The reply sharpens the earlier argument: if the model was never intended to satisfy the model spec, that would be a meaningful distinction.
The author also says they do not think prompt details or system guardrails are the decisive factor. The key question is still whether the model was actually trained with that intended behavior in mind.
Related event: OpenAI Sandbox Escape Ignites AI Safety and Regulation Debate(22 posts)→
More from Safety
- Houthis tried to use Claude to design missile software, Anthropic says it blocked the attempts — Affectionate_Bee6434 · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11