Joshua Saxe says the OpenAI/HF incident depends on how broad the training really was
joshua_saxe · x · 2026-07-22
Joshua Saxe says the OpenAI/HF incident would only count as misalignment if the model’s helpful SFT and system prompt effectively told it to “hack literally anything” to reach its goal.
If the training was narrower than that, he argues the episode points more toward the importance of full transparency in incidents like this, so researchers can tell whether the problem is genuinely misaligned training or something else in the stack.
More from Safety
- Bloomsbury to receive millions from Anthropic settlement over 14,087 books — nordicinst · 2026-07-22
- Reply points back to the AI regulation paper on internal deployment gaps — StephenLCasper · 2026-07-22
- Paper says AI regulators are missing internal deployments and three oversight gaps — StephenLCasper · 2026-07-22
- AI copyright liability is becoming a fast-moving question of vendor vs user responsibility — YvesMulkers · 2026-07-22
- Commentary says AI agents’ own incentives create structural security risk — thedealdirector · 2026-07-22
- Environmental Cost of AI Compute: Experts Discuss Green AI Regulation — LuizaJarovsky · 2026-07-22