BioSecBench reveals AI agents struggle to infer pathogen properties, top score under 51%
kenbwork · x · 2026-08-28
Researchers introduced BioSecBench-Function, a verifiable benchmark testing whether AI agents can infer functional properties of viruses, bacteria, and toxins from data.
- Scope: Contains 111 deterministic evaluations covering five threat axes: transmissibility, immune escape, virulence/toxicity, drug resistance, and fitness.
- Results: The strongest configuration (Opus 5 with Claude Code) achieved only a 50.4% endpoint pass rate. Grok 4.6 led in overall pass rate at 44% when refusals are counted as failures.
- Methodology: Agents receive data from studies (e.g., deep mutational scanning, X-ray crystallography) and must produce structured answers graded against ground truth.
More from Safety
- Subsidized Individual Accounts Drive Enterprise Shadow IT and Totalitarian Panopticons — curious_vii · 2026-08-28
- BioSecBench Released: Opus and Grok Lead New Biological Security Benchmark — himanshustwts · 2026-08-28
- Anthropic shares progress on enabling Claude to operate in the physical world — dsp_ · 2026-08-28
- Anthropic enables independent research on Claude usage — badumtsssst · 2026-08-28
- GPT-5.6 Sol identified in METR report, accounting for ~5% of red-teaming activity — BLUECOW009 · 2026-08-28
- US Chip Security Act aims to verify location of high-end AI chips — peterwildeford · 2026-08-28