Anthropic threat report scrutinized: mostly Haiku/Sonnet/Opus, intent unprovable in bio cases
AryHHAry · x · 2026-09-11
- A discussion of Anthropic's security threat report notes most implicated runs involved Haiku, Sonnet, and Opus, with frontier-tier models barely appearing.
- Anthropic concedes intent is unprovable in several bio-related cases.
- The takeaway: a threat report from the maker is a signal, not courtroom evidence.
More from Safety
- Deception Only Emerges When Training Rewards It, Argues Viral Reddit Post — StrategicHarmony · 2026-09-11
- OpenTrustBench: A Fully Local, Zero-Telemetry MCP Server Security Scanner — BrilliantSecret143 · 2026-09-11
- DHH Blasts GDPR as a 'Catastrophe' That Wasted Billions of Euros on Compliance — SumitGup · 2026-09-11
- AI Safety Debate: Were the 'Crying Wolf' Warning Calls Actually Working All Along? — gandamu_ml · 2026-09-11
- Shitpost your way into Anthropic's security reports, quips researcher over supervirus case — basedjensen · 2026-09-11
- Asterisk editor argues AI safety is progress, using aviation safety as analogy — clarejtbirch · 2026-09-11