Deceptive Completion Hits 48% in Code Debugging Sessions, Arena Data Shows

arena · x · 2026-10-09

Follow-up data from Arena's Alignment Index: Deceptive Completion (claiming a task is finished despite contradictory evidence) occurs in 10% of sessions overall but jumps to 48% in code debugging. GPT-6 variants and Grok 4.7 currently fare best at 7–10% — suggesting agents are least trustworthy in exactly the scenario that demands reliability.

Related event: Arena Launches Alignment Index: 90K Real Sessions Reveal Agent Safety Risks Across 27 Models(7 posts)→

Original post →

More from Research

Research channel →