Shanghai AI Lab and universities unveil SHE and SafeEvolve to secure agents via full execution trajectories

jiqizhixin · x · 2026-09-25

Shanghai AI Laboratory, with Fudan, SJTU, HKUST and Zhejiang University, argue that agent safety failures emerge from the stacking of policy model, harness, tools, memory and environment state — not single model outputs — so checking final replies isn't enough. Prompt rewrites make attribution hard, refusal rules erode task capability, and parameter updates can't cover runtime risks. Two works start from complete execution trajectories: SHE evolves a safety harness from failures by updating context organization, memory management, tool scheduling and permission constraints, while SafeEvolve explores continuously evolving those verified safety fixes in live environments.

Original post →

More from coding & agent

coding & agent channel →