SeerGuard uses a world model to screen risky actions in mobile GUI agents
Xue Yu · hf · 2026-07-22
- The paper proposes SeerGuard, a consequence-aware safety framework for mobile GUI agents.
- It combines pre-execution instruction screening with action-level risk assessment so the system can predict likely outcomes before the agent acts.
- The authors build a unified safety-augmented world model (SAWM) via multi-task learning, mixing next-state prediction and safety risk assessment.
- On Qwen3-VL-8B-Instruct, the safety-utility score rises from 0.191 to 0.596 at ω=0.8, while risk-cost drops from 0.347 to 0.130 at α=0.8.
- The paper argues the model generalizes across diverse mobile GUI agents and shows both screening and action-risk prediction contribute to the gains.
More from coding & agent
- Claude Code adds iOS Simulator control for side-by-side mobile testing — xiaohu · 2026-07-22
- Salesforce says enterprise agents are moving from loops to graph routing — msrivastav13 · 2026-07-22
- Builders ask how to test PayPal and Cash App flows for AI agents without real identities — Friendly-Tear990 · 2026-07-22
- You should default to small subagent teams for well-scoped tasks — alex_teichman · 2026-07-22
- Cursor team is building Cursor with Cursor, inside Cursor — soleio · 2026-07-22
- NVIDIA shows Unreal Engine wired to Claude Code and Cursor via MCP — nptacek · 2026-07-22