DeepMind Paper: One Misleading Hint Drops Coding Agent Scores by Up to 46.7%

dair_ai · x · 2026-09-25

A new Google DeepMind paper introduces XYEval, which adds one confident but wrong hint to tasks from tau2-bench, SWE-bench, Terminal-Bench, HLE and MCP-Atlas.

Original post →

More from coding & agent

coding & agent channel →