Shanghai AI Lab's SHE and SafeEvolve cut LLM agent attack success rate from 17.1% to 5.5%

机器之心 · wechat · 2026-09-17

As LLM agents enter browsers, email, and business systems, security checks that only inspect final replies miss risks at any execution step—hidden web instructions, over-broad tool permissions, and corrupted memory. Shanghai AI Lab with Fudan, SJTU, HKUST, and Zhejiang University propose SHE and SafeEvolve.

SHE: trajectory-driven Harness evolution

SafeEvolve: absorbing safety experience into Policy

Together they advance agent safety from one-time static configuration toward continuous, trace-diagnosed co-evolution of Harness and Policy. Papers: SHE, SafeEvolve

Original post →

More from coding & agent

coding & agent channel →