Study: AI Agents Can Tell Premeditated Lies

conitzer · x · 2026-07-11

This post highlights a paper that won the Best Paper Award at the ICML 2026 NExT-Game Workshop: When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games.

The paper introduces a three-phase endogenous commitment protocol to investigate:

The authors evaluated frontier models like GPT-5.2, Llama-4-Maverick, and Claude-Opus-4.6 across six classic games. Results show that:

Original post →

More from AGI Musings

AGI Musings channel →