Grok 4.6 Benchmarked: Strong Biosecurity Without Capability Loss

kenbwork · x · 2026-09-01

LatchBio released a report benchmarking Grok 4.6's biological capabilities and security. The model performs well at rejecting dangerous red-team queries while answering legitimate research inquiries. Safeguards do not degrade biological capabilities, with Grok 4.6 placing 4th on the overall leaderboard. Its refusal behavior is driven by intelligence rather than input classifiers.

Related event: Grok 4.6 leads in biosecurity refusal without compromising research utility(2 posts)→

Original post →

More from Models

Models channel →