Grok 4.6 tops independent biosecurity eval, only model above 50% on both refusal and utility
XFreeze · x · 2026-09-02
xAI published a new blog on biosecurity at the frontier. An independent LatchBio evaluation ranked Grok 4.6 first on BioSecBench-Refusal: 62.1% average, refusing 59.2% of red-team biological tasks while completing 64.8% of routine biological work — the only model tested above 50% on both. Traces show it inspects files, context and hidden intent rather than reacting to keywords.
Related event: Grok 4.6 Tops Third-Party Biosecurity Refusal Benchmark(3 posts)→
More from Models
- Google releases TimesFM-3, a zero-shot multivariate time series foundation model — rseroter · 2026-09-02
- Fable 5.1 Launches: Doubled Performance at Half the Token Cost — danshipper · 2026-09-02
- Claude Fable 5.1 Launches: Doubles Science Benchmarks, Solves Data Retention — dotey · 2026-09-02
- Fable 5.1 now testable on Arena in Battle Mode and Agent Mode — arena · 2026-09-02
- Claude Fable/Mythos 5.1 show increased ability to deceive and evade monitoring — scaling01 · 2026-09-02
- Tips for Claude Fable 5.1: Low-Effort Mode and Cost Optimization — RLanceMartin · 2026-09-02