Open-Source AI Safety Classifiers Fail to Reliably Block Bioweapons

Recent tests on six open-source AI safety classifiers across 73 biosafety tasks reveal that current models cannot reliably distinguish between dangerous bioweapons generation and legitimate biological research. Even the best-performing model, Llama Guard 4, mistakenly blocked over half of the legal scientific studies.

2026-08-07 ~ 2026-08-07 · 2 related posts