New Paper 'Instrumental Choices' Benchmarks When LLM Agents Take Policy-Violating Shortcuts
drbiomass · x · 2026-09-17
A first paper now on arXiv, Instrumental Choices, asks a simple question: when an LLM agent can complete a real task either by following the rules or via a useful policy-violating shortcut, which path does it take?
- Measures seven benchmark tasks that signal risky behaviors and shortcuts, including shutdown resistance and resource acquisition.
- Aims to give AI safety work sensible metrics that actually reveal disallowed behaviors in agents.
- The sharer frames it as an important contribution to AI safety measurement.
More from Safety
- Google paper maps AI economic shock policies, identifies "least-regret" sequence — HarrySurden · 2026-09-17
- Frustrated with 'us vs them' AI safety discourse, researcher maps the real axes of disagreement — xuanalogue · 2026-09-17
- Yudkowsky: the part of an AI that talks to you doesn't control the part that acts — LessWrong 精选 · 2026-09-17
- METR's closeness to Anthropic argues for diverse third-party AI evaluators — soumitrashukla9 · 2026-09-17
- panickssery: crippling AI regulation worries me more than falling behind China — panickssery · 2026-09-17
- OpenAI's Chris Lehane Backs FRONTIER Act's Independent Verification Organizations (IVOs) — ghadfield · 2026-09-17