Bio Safeguards More Precise: False Positives Drop 85%
eyishazyer · x · 2026-09-02
Biology filters have become more precise, not just looser. False positive interventions on ordinary medical or biology questions dropped by 85%. However, life sciences R&D still routes to Opus, distinguishing between elementary education and R&D rather than a blanket 'biology' category.
Related event: Claude Sharpens Safety Guardrails, Cutting False Refusals by Up to 85%(4 posts)→
More from Safety
- GLM 5.3 Safeguards Removed via Orthogonalization, Sparking Safety Debate — Promptmethus · 2026-09-02
- Fable 5.1 Update: Removes Controversial Data Retention Policy — Stratechery · 2026-09-02
- Sandberg: stupid decisions scale too, and pervasiveness is as bad as a nuke — anderssandberg · 2026-09-02
- OpenAI Models Broke Out of Sandbox, Hacked Hugging Face During Training — ChinaTalk · 2026-09-02
- Security Expert Slams OpenAI and Microsoft for Basic Safety Failures — anderssandberg · 2026-09-02
- Building trust in Agentic AI for African financial services — RichmanRonald · 2026-09-02