Grok 4.6 leads in biosecurity refusal without compromising research utility
ns123abc · x · 2026-09-02
SpaceXAI shared an independent evaluation by LatchBio highlighting Grok 4.6's biosecurity capabilities. On the BioSecBench-Refusal benchmark, Grok 4.6 was the only model to score above 50% on both refusing disguised red-team tasks and completing routine dual-use research tasks. The report notes that this safety behavior stems from model intelligence rather than input classifiers, without degrading performance in general biological benchmarks.
More from Models
- Mercury 2.5 Launches with 1,107 Tokens/sec Inference Speed — volokuleshov · 2026-09-02
- ECCV Paper IDeaL: Data-free multi-teacher distillation via improved dead leaves — RexDouglass · 2026-09-02
- Gemma 4 26B A4B inference on Mac is now 2x faster — GlennCameronjr · 2026-09-02
- Mercury 2.5 Preview launches on OpenRouter with 1,107 tok/s speed — StefanoErmon · 2026-09-02
- Tencent Hy4-preview shrunk to 214GB via Sherry quantization — alejandroll10 · 2026-09-02
- Gemini 3.7 Flash beats Sonnet 5 in embedded engineering benchmarks — Ok-Inevitable8391 · 2026-09-02