Grok 4.6 leads spatial biology benchmark but fails biosecurity tests
kenbwork · x · 2026-08-15
Citing benchmark data, kenbwork notes that Grok 4.6 (Grok Build) achieves an impressive 74.8% on the spatial biology benchmark. However, it ranks as the worst model tested for biosecurity, scoring only 1.6% on the refusal benchmark, indicating a significant lack of safety guardrails in this context.
More from Models
- Tested 3 models to spec a local AI-brain install: one cited real files, one got macOS compat backwards — schwentker · 2026-08-15
- Questioning how to detect Gemini degradation — ngxson · 2026-08-15
- User Feedback: Opus 5 Seems Too Histrionic — curious_vii · 2026-08-15
- OpenAI's Unreleased 'Astra' Model Solves 10 Math Problems — thursdai_pod · 2026-08-15
- Long Reasoning Task Consumes 22k Tokens for 3k Output — generativist · 2026-08-15
- Users Report Codex Asking for Permission More Often — GabGarrett · 2026-08-15