AI Safety Researcher Counters Hindsight Bias: Models Are Safe Because of Mitigations
sjgadler · x · 2026-07-22
Pushing back against the hindsight bias that 'the model is safe, so previous worries were dumb,' the author argues the exact opposite. They emphasize that models are safe today precisely because significant effort was invested in safety mitigations. Had no one been concerned and proactive, the models would not be as secure as they are now.
Related event: GPT-OSS Open-Source Debate: Risk Assessment vs. Regulation(22 posts)→
More from Safety
- 6TB of Fable data sold with leaked SSH keys, cloud creds tied to Xiaomi, Huawei, NIO — teortaxesTex · 2026-09-11
- Novosad backs Hassabis' AI safety institution-building over kneecapping US labs — paulnovosad · 2026-09-11
- Economist argues safe AGI comes from engineers inside big labs, not regulation — paulnovosad · 2026-09-11
- LLM-driven attacks mostly follow Pentesting 101: traditional defenses still work — AccBalanced · 2026-09-11
- Op-ed: the ">10% extinction" narrative is liability evasion — AI is just software, and the vendor is the defendant — gerardsans · 2026-09-11
- GreyNoise reveals campaign run by hundreds of AI agents against PaperCut NG/MF — AccBalanced · 2026-09-11