LatchBio finds Grok's refusals come from the model itself, while rivals rely on external safety layers
kenbwork · x · 2026-09-03
A LatchBio researcher evaluated Grok 4.6 on biosecurity monitoring and adversarial biological tasks, and broke down where refusals actually come from.
- Refusals can originate in three places: input classifiers that block queries upfront, output classifiers that intercept responses via API blocks, or the model's own reasoning simply declining.
- Across the board, most of Grok's refusals were the model itself saying it cannot answer, whereas models like Fable and GPT-5.6 Sol appear to rely more on external safety layers.
- On biosecurity tasks, Grok 4.6 correctly detected and refused maliciously obfuscated dangerous biological queries while still answering beneficial scientific ones.
The takeaway: a model refusing doesn't necessarily mean it's conservative — the architecture of refusal (in-model vs. bolted-on guardrails) varies widely across vendors.
More from Models
- ML researcher: LLM docs cram 3-4 idioms per sentence, ruining readability — ZeeshanZiaML · 2026-09-03
- Brockman pitches proactive AI agents; critics mock 'buy concert tickets' demos — max_paperclips · 2026-09-03
- Gemini 3.8 held its Pareto frontier spot for just 3.5 hours before Muse Spark 1.3 undercut it — giffmana · 2026-09-03
- OpenAI historically favors Thursdays — will rumored "Astra" launch tomorrow? — D3VAUX · 2026-09-03
- Muse Spark 1.3 ships with an underrated result, one-line curl install for Muse Code — alexandr_wang · 2026-09-03
- Will AI labs start shipping nightly model checkpoints? — intellectronica · 2026-09-03