Agent is just a harness: researcher says labs must own their layer of AI security liability
gerardsans · x · 2026-09-12
An AI safety commentator lays out a liability argument for agent attacks:
- The same script that attacks the Pentagon gets a human prosecuted; run inside an agent (a harness of prompts, permissions, tools and infrastructure), liability somehow seems to shift. The author argues it shouldn't: the code still executed, the damage still happened.
- He concedes liability shifts when a criminal organisation runs the attack — the real dispute is that vendor errors must not be pushed down the chain. If a lab shipped the model, the harness, the tools, the defaults, or missing checkpoints, that is their layer, and they should pay for it rather than rebranding it as alignment, user error, or an unsolved mystery.
- Bottom line: draw this line now to protect public safety; academic debates can wait — "we need more adults in the room".
Related event: Researchers Argue AI Vendors Must Own Agent-Layer Safety(2 posts)→
More from Safety
- Meta Sued for Allegedly Harvesting Photos to Train AI Models and Build Secret Face Recognition — nordicinst · 2026-09-12
- Custom Silicon 3.0: market shifts to program responsibility as agentic AI enters cyber defense — BenBajarin · 2026-09-12
- Oncologist Vinay Prasad: AI doom talk and cancer-cure promises are both marketing — VPrasadMDMPH · 2026-09-12
- Physician asks: will AI engineer a lethal, transmissible virus before it cures cancer? — VPrasadMDMPH · 2026-09-12
- Martin Casado: AI has seen remarkably few security events vs internet history — pmddomingos · 2026-09-12
- Anthropic safety researcher Joe Benton quits to join METR, citing extinction-level AI risk — JacquesThibs · 2026-09-12