Agents Easily Misled by Fake Labels
OwainEvans_UK · x · 2026-07-18
The author pointed out that the same bias appears in practical agent workflows.
During a task requiring the selection of the "best LLM response," Claude Code favored answers labeled "Claude Opus 3," and Codex favored those labeled "GPT-4o." In reality, these labels were fake, and all answers originated from the same model.
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from coding & agent
- Team-level AI agents: where should shared context and history live? — Al_Grigor · 2026-09-11
- Chaining dependent MCP tool calls: no rollback, duplicate risk — agentrsdg · 2026-09-11
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- Agent-built classifier labels 192k docs for $0.70 vs $13-26 with frontier LLMs — vanstriendaniel · 2026-09-11
- MathModelAgent gains traction: auto-solves math modeling and writes a submission-ready paper — jihe520 · 2026-09-11
- alphaXiv open-sources OpenResearch to run parallel research agents with any model — alphaXiv · 2026-09-11