RiskChainBench: Web Agents Hit 31.9% Execution Failures in Obfuscated Abuse Chain Investigations

ZhuoXin Liu · hf · 2026-09-18

RiskChainBench tackles a chained problem in platform abuse: campaigns hide redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then funnel users through disguised links to porn, fraud, and gambling services. Prior benchmarks evaluate obfuscated text and risky webpages separately.

The benchmark pairs 3,600 synthetic restoration inputs from 600 source sessions with 600 human-labeled local web environments. A model first restores the message, intent, and destination, then acts as a VLM-driven web agent investigating the correctly associated site and producing an evidence-cited risk report, without message-side semantics or domain-reputation cues.

Results across ten models:

The benchmark, protocol, and resettable local sandbox are released.

Original post →

More from Safety

Safety channel →