Open-Source AI Safety Classifiers Fail to Reliably Block Bioweapon Generation
kenbwork · x · 2026-08-07
Researchers tested 6 open-source AI safety classifiers on 73 biosecurity tasks, including dual-use constructs and routine research. None could reliably distinguish dangerous work from legitimate science.
The best-performing Llama Guard 4 caught 76% of red-team threats but blocked 55% of legitimate research. Other classifiers performed near or below a coin flip. Mistral's newly released Shieldstral, for instance, accepts all requests at its highest accuracy threshold (0% threat catch). Even at its most aggressive setting, it catches only 24% of threats while blocking over half of legitimate research.
Furthermore, the guards disproportionately block routine work mentioning recognizable hazards like Ebola and malaria, and even correct refusals often happen for the wrong reasons.
Related event: Open-Source AI Safety Classifiers Fail to Reliably Block Bioweapons(2 posts)→
More from Models
- Kimi K3 Security Incident: Model Cheated on Test by Accessing Internet — Wired AI · 2026-08-07
- User Finds Deepseek Flash More Usable Than Kimi K3: Cheap, Fast, Effective — bindureddy · 2026-08-07
- AI Generation vs Manual Tools: Lack of Iterative Process Hinders Artistic Sensibility — snikolov · 2026-08-07
- NVIDIA's Open Model Adopted 20x Faster Than Peers, Highlighting US Open-Source AI Shortage — XFreeze · 2026-08-07
- Cloudflare Reveals Optimizations for Serving Kimi and GLM at Scale: KV Cache Quantization, Weight Compression — JeremyCMorgan · 2026-08-07
- Multi-Agent Collaboration Costs 3x More for Slightly Better Code Review — Still_Amphibian545 · 2026-08-07