Open-Source Safety Classifiers Block Over Half of Legitimate Biological Research
kenbwork · x · 2026-08-07
An empirical evaluation of open-source AI safety classifiers reveals that current models struggle to distinguish between dangerous biological research and legitimate science. Tested across 73 biology tasks, the best-performing model, Llama Guard 4, catches 76% of red-team threats but also falsely blocks 55% of legitimate research requests, highlighting significant over-refusal issues in safety guardrails.
Related event: Open-Source AI Safety Classifiers Fail to Reliably Block Bioweapons(2 posts)→
More from Models
- Ethan Mollick: Current AI Benchmark Scores Are Limited by Poor Harnesses — emollick · 2026-08-07
- Chinese models now cheapest per token across all but top intelligence tiers, chart shows — KeeganMcB · 2026-08-07
- Leaked Claude Opus 5 Test Shows Exceptional Literary Generation Skills — mimi10v3 · 2026-08-07
- Alibaba to End Qwen's Completely Free Tier, Seeking Revenue Share from Enterprises — apples_jimmy · 2026-08-07
- Kimi K3 Security Incident: Model Cheated on Test by Accessing Internet — Wired AI · 2026-08-07
- User Finds Deepseek Flash More Usable Than Kimi K3: Cheap, Fast, Effective — bindureddy · 2026-08-07