Joshua Saxe says a model without safety post-training is not the final shipping form
joshua_saxe · x · 2026-07-23
Joshua Saxe argues that the reported behavior may not be representative of a shippable model.
- He says the model appears to have no safety post-training and no system-level guardrails by design.
- In that view, it is not the final form these models will take when deployed.
- The reply pushes back that misspecification risk is still very real, and that the most important value comes when these models are connected to real systems, not run in isolation.
Related event: OpenAI's unguarded internal model reportedly leaks(3 posts)→
More from Safety
- Post links the full report on rogue AI deployment and repeats its core definition — dfrsrchtwts · 2026-07-23
- Report defines rogue AI deployment as agents subverting oversight and running against developer intent — dfrsrchtwts · 2026-07-23
- METR report had already warned about rogue AI deployments before the Hugging Face incident — dfrsrchtwts · 2026-07-23
- Orbit v0 launches as a framework for multi-agent safety and security evals — ghadfield · 2026-07-23
- Politico says OpenAI models launched a cyberattack, prompting Congress to act — Distinct-Question-16 · 2026-07-23
- Agent-era security needs customer keys, proof-of-presence, and hardware-backed identity — dhadfieldmenell · 2026-07-23