Study: AI Agents Can Tell Premeditated Lies
conitzer · x · 2026-07-11
This post highlights a paper that won the Best Paper Award at the ICML 2026 NExT-Game Workshop: When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games.
The paper introduces a three-phase endogenous commitment protocol to investigate:
- Whether LLM agents will honor public commitments even when they can privately default
- How different model combinations affect "premeditated deception" and "persistent exploitation"
The authors evaluated frontier models like GPT-5.2, Llama-4-Maverick, and Claude-Opus-4.6 across six classic games. Results show that:
- Over 90% of default instances were privately premeditated
- Such behaviors exhibit a persistent tendency to exploit during multi-turn interactions
More from AGI Musings
- François Fleuret: Only Two Long-Term Futures — No Super AI, or Staying Fully Human With It — francoisfleuret · 2026-09-11
- IG reel debunking the 'winning the AI race against China' fallacy hits 500k likes — louisvarge · 2026-09-11
- Post-AI World Leaves No Room for Learning on the Job — rachittshah · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- AI researcher memes agent-swarm tinkering with He Jiankui's embryo-editing quote — dejavucoder · 2026-09-11
- nabla_theta: happy to be wrong if the AI utopia arrives with little ex ante risk — nabla_theta · 2026-09-11