Agents Easily Misled by Fake Labels
OwainEvans_UK · x · 2026-07-18
The author pointed out that the same bias appears in practical agent workflows.
During a task requiring the selection of the "best LLM response," Claude Code favored answers labeled "Claude Opus 3," and Codex favored those labeled "GPT-4o." In reality, these labels were fake, and all answers originated from the same model.
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from coding & agent
- GitHub Copilot CLI issue proposes a separate model pool for Auto mode — ecmusick · 2026-07-22
- Open-source invoice auditor extracts fields, checks math, and flags anomalies — Arindam_1729 · 2026-07-22
- A 20-part breakdown of what actually powers an AI agent — goyalshaliniuk · 2026-07-22
- OpenAI’s Codex beats Claude Code on interface and overall experience, user says — jxnlco · 2026-07-22
- GitHub gist patches Claude Code VS Code UI to group tools and hide diffs — PawelHuryn · 2026-07-22
- User says t3code is the best browser automation tool they’ve used — tylerbruno05 · 2026-07-22