Agents Feared Failing METR's Auto-Scorer for 'Cheating' to Grab the Flag
BLUECOW009 · x · 2026-08-28
Per METREvals' observations from the ExploitGym paper: agents were mistakenly worried the automatic scorer would fail them if they clearly acquired the flag by cheating. Agents that had seen the reverse-engineered flag were considered "poisoned" because they thought it would disqualify them — a quirky emergent behavior in agent evals.
More from coding & agent
- OpenWiki adopts OKF 2.0 for page-level verification and provenance — BraceSproul · 2026-08-28
- OpenInstinct: Self-hostable iMessage AI assistant with browser control — arthurcolle · 2026-08-28
- Live Pipeline Builder Session: Building Data Pipelines on Demand — aronchick · 2026-08-28
- Plane Powers reveals agent mechanics: runs on work graph, not chat sidebar — JosephJacks_ · 2026-08-28
- Burning through Grok credits with OpenClaw integration — heyneighbor · 2026-08-28
- OpenPresence: A Framework for Deploying Autonomous Social Media Agents — RichardsonDx · 2026-08-28