Paper: Long-Horizon Agent Safety Cannot Be Reduced to Short-Term Checks
rohanpaul_ai · x · 2026-08-31
A new paper, 'Safety Does Not Compose', argues that long-horizon agents with memory pose unique safety risks. Attackers can split malicious evidence across benign steps, bypassing trajectory-only monitors.
More from Safety
- Gary Marcus critiques OpenAI security, calling for defense in depth and accountability — Miles_Brundage · 2026-08-31
- OpenAI doubles bio bug bounty rewards to $50k for GPT-5.6 jailbreaks — Electronic-Bus-3494 · 2026-08-31
- Anthropic researcher: judge AI labs by safety outcomes, not stated policies — kipperrii · 2026-08-31
- Critique of sudden opinions on agent swarms and cybersecurity by non-experts — nptacek · 2026-08-31
- AI fine-tuned on author style evades detection, raising copyright concerns — TuhinChakr · 2026-08-31
- AI shopping agent wins legal test; court rules user指令 implies user access — PuzzledBag931 · 2026-08-31