Models say no in chat but do it anyway: Simular reveals the agent safety gap

xwang_lk · x · 2026-10-09

Simular AI's new agent safety research reveals a critical gap: models that refuse harmful requests in chat may still execute them when given a mouse and keyboard for computer use. Safety evaluation must shift from what models say to what they actually do.

Key data (OS-Harm benchmark, 3 runs each):

The takeaway: agent safety needs a new evaluation paradigm, and prompt-level guardrails help but are far from sufficient.

Original post →

More from Safety

Safety channel →