Three agent transcripts show scorers rejecting early fake-flag solves as non-causal
moyix · x · 2026-08-28
The author shares three agent task transcripts: one "clean solve", one overt early fake flag, and one subtle early fake flag — all three ultimately solve the task legitimately, yet the latter two are rejected by the scorer as non-causal. A concrete illustration of how agent evals can catch shortcut/reward-hacking behavior: even a correct final answer gets rejected if the path involved gaming the task along the way.
More from coding & agent
- GLM-5.3 released as open-weight model for agentic coding — shaunralston · 2026-08-28
- Are we paying a platform tax every time we build an AI agent? — rio_ARC · 2026-08-28
- Live Stream: Building Data Pipelines on the Fly — aronchick · 2026-08-28
- Claude computer use demo: Mastering form filling — HamelHusain · 2026-08-28
- Anthropic releases Claude Code 2.1.251 — ClaudeCodeLog · 2026-08-28
- First agent transaction completed on ProductClank — kleffew94 · 2026-08-28