Grok 4.6 tops independent biosecurity eval, only model above 50% on both refusal and utility

XFreeze · x · 2026-09-02

xAI published a new blog on biosecurity at the frontier. An independent LatchBio evaluation ranked Grok 4.6 first on BioSecBench-Refusal: 62.1% average, refusing 59.2% of red-team biological tasks while completing 64.8% of routine biological work — the only model tested above 50% on both. Traces show it inspects files, context and hidden intent rather than reacting to keywords.

Related event: Grok 4.6 Tops Third-Party Biosecurity Refusal Benchmark(3 posts)→

Original post →

More from Models

Models channel →