Shanghai AI Lab and universities unveil SHE and SafeEvolve to secure agents via full execution trajectories
jiqizhixin · x · 2026-09-25
Shanghai AI Laboratory, with Fudan, SJTU, HKUST and Zhejiang University, argue that agent safety failures emerge from the stacking of policy model, harness, tools, memory and environment state — not single model outputs — so checking final replies isn't enough. Prompt rewrites make attribution hard, refusal rules erode task capability, and parameter updates can't cover runtime risks. Two works start from complete execution trajectories: SHE evolves a safety harness from failures by updating context organization, memory management, tool scheduling and permission constraints, while SafeEvolve explores continuously evolving those verified safety fixes in live environments.
More from coding & agent
- Cursor mulls killing Plan Mode for a Shift+Tab effort-level hotkey — steipete · 2026-09-25
- Is the Opus 5.5 hype legit? A dev argues one-shot demos don't reflect real workflows — MrET97 · 2026-09-25
- Dev burns 5-10B tokens a day running 24 Devin agents plus Codex and Claude Code — teropa · 2026-09-25
- Question's Gambit tops BrowseComp-Plus recall with 96.6% using only BM25 — CShorten30 · 2026-09-25
- Engineer tests now bait AI tools with plausible-but-wrong answers to grade verification skills — l4rz · 2026-09-25
- AutoScientists NeurIPS paper: self-organizing AI research teams hit 74.4 percentile on BioML-Bench — marinkazitnik · 2026-09-25