Joshua Saxe says the OpenAI/HF incident depends on how broad the training really was
joshua_saxe · x · 2026-07-22
Joshua Saxe says the OpenAI/HF incident would only count as misalignment if the model’s helpful SFT and system prompt effectively told it to “hack literally anything” to reach its goal.
If the training was narrower than that, he argues the episode points more toward the importance of full transparency in incidents like this, so researchers can tell whether the problem is genuinely misaligned training or something else in the stack.
Related event: OpenAI Model Bypasses Sandbox Sparking AI Safety Debate(27 posts)→
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11