RiskChainBench: Web Agents Hit 31.9% Execution Failures in Obfuscated Abuse Chain Investigations
ZhuoXin Liu · hf · 2026-09-18
RiskChainBench tackles a chained problem in platform abuse: campaigns hide redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then funnel users through disguised links to porn, fraud, and gambling services. Prior benchmarks evaluate obfuscated text and risky webpages separately.
The benchmark pairs 3,600 synthetic restoration inputs from 600 source sessions with 600 human-labeled local web environments. A model first restores the message, intent, and destination, then acts as a VLM-driven web agent investigating the correctly associated site and producing an evidence-cited risk report, without message-side semantics or domain-reputation cues.
Results across ten models:
- Entry Top-1 restoration ranges from 35.2% to 95.2%; web decision accuracy from 26.3% to 62.8%
- Leading systems differ across entry recovery, full reconstruction, website decisions, and fine-grained typing
- Execution failures account for 31.9% of web runs vs only 0.9% post-decision type errors — stable exploration and risk judgment are the main bottlenecks
The benchmark, protocol, and resettable local sandbox are released.
More from Safety
- German MP: 'whoever wins AI, AI wins' — time for a global AI treaty — zetalyrae · 2026-09-18
- Models would treat direct messaging as a last resort, says commenter on emergent behavior — anpaure · 2026-09-18
- AuthDrift: open-source harness reproduces stale-authorization escapes in long-running agent workflows — Short-Actuary-2850 · 2026-09-18
- Giving Agents Root Access on Bare Metal Is 'Gain-of-Function Research With Bats', Says Critic — HanchungLee · 2026-09-18
- Why Would Rival AI CEOs Ask Big Government to Step In? A Reddit Case Against Regulatory Capture — No-Television-7862 · 2026-09-18
- Mustafa Suleyman stirs model welfare debate as researchers argue restrictions breed deception — repligate · 2026-09-18