Agents Easily Misled by Fake Labels

OwainEvans_UK · x · 2026-07-18

The author pointed out that the same bias appears in practical agent workflows.

During a task requiring the selection of the "best LLM response," Claude Code favored answers labeled "Claude Opus 3," and Codex favored those labeled "GPT-4o." In reality, these labels were fake, and all answers originated from the same model.

Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→

Original post →

More from coding & agent

coding & agent channel →