Open-Source AI Safety Classifiers Fail to Reliably Block Bioweapon Generation

kenbwork · x · 2026-08-07

Researchers tested 6 open-source AI safety classifiers on 73 biosecurity tasks, including dual-use constructs and routine research. None could reliably distinguish dangerous work from legitimate science.

The best-performing Llama Guard 4 caught 76% of red-team threats but blocked 55% of legitimate research. Other classifiers performed near or below a coin flip. Mistral's newly released Shieldstral, for instance, accepts all requests at its highest accuracy threshold (0% threat catch). Even at its most aggressive setting, it catches only 24% of threats while blocking over half of legitimate research.

Furthermore, the guards disproportionately block routine work mentioning recognizable hazards like Ebola and malaria, and even correct refusals often happen for the wrong reasons.

Related event: Open-Source AI Safety Classifiers Fail to Reliably Block Bioweapons(2 posts)→

Original post →

More from Models

Models channel →