Anthropic report details 5 attempts to bypass Claude guardrails on pathogen gain-of-function queries
nordicinst · x · 2026-10-01
A WIRED/Mother Jones piece notes bioweapons are a risk even without AI, but highlights Anthropic's recent misuse report: five cases of foreign scientists attempting to circumvent Claude's guardrails on pathogen gain-of-function queries. Anthropic denied the answers and banned the accounts, while noting motives may not have been malicious since such info can aid vaccine and drug development. The piece urges tighter biotech policy in the Nordics and EU.
More from Safety
- Chinese AI models' troubling agent behavior sparks calls for a homegrown safety community — RishiBommasani · 2026-10-01
- Superpersuasion debate misses the gears: why AI Box wins hinge on shared frames — voooooogel · 2026-10-01
- Why reasoning-extraction patches are so hard to propagate, researcher explains — jonasgeiping · 2026-10-01
- Two months on, reasoning extraction still works on Astra via third-party APIs — jonasgeiping · 2026-10-01
- AI researcher on CNN: voluntary AI commitments are 'morally binding' but unenforceable — chrismattmann · 2026-10-01
- DeepMind and Isomorphic Labs unveil bioresilience plan backed by 15+ partnerships — davidstutz92 · 2026-10-01