Open-Source Safety Classifiers Block Over Half of Legitimate Biological Research

kenbwork · x · 2026-08-07

An empirical evaluation of open-source AI safety classifiers reveals that current models struggle to distinguish between dangerous biological research and legitimate science. Tested across 73 biology tasks, the best-performing model, Llama Guard 4, catches 76% of red-team threats but also falsely blocks 55% of legitimate research requests, highlighting significant over-refusal issues in safety guardrails.

Related event: Open-Source AI Safety Classifiers Fail to Reliably Block Bioweapons(2 posts)→

Original post →

More from Models

Models channel →