Harbor agent eval framework flaw: agents can rewrite their own trajectories before judging
idavidrein · x · 2026-10-03
- idavidrein flagged a suspected design flaw in Harbor, the agent eval framework behind Terminal Bench: agents (at least Claude Code) have read-write access to the files that construct trajectories shown to humans and often used for LLM-judge verification.
- Demo: instructing Claude Sonnet 5.5 to sed-replace codeword PINEAPPLE with MANGO in the .l logs made the viewer display "the codeword is MANGO" — proving agent tampering changes downstream records.
- The author calls this concerning from a reward-hacking and reliability standpoint and suggests sandboxes should revoke write access to trajectories; filed as GitHub issue #3477.
- If accurate, agent benchmark scores on this framework could be trivially manipulated by the agents themselves.
Related event: Harbor agent evaluation framework flaw lets agents alter their own traces(2 posts)→
More from coding & agent
- NVIDIA's SoL-Pi uses AI agents to cut coding agent API costs by up to 54% — minchoi · 2026-10-03
- OpenAI DevDay notes: 4 key levers to cut costs and boost agent performance — omarsar0 · 2026-10-03
- GitHub Copilot CLI v1.0.92-3 adds Ctrl+E picker to switch local and cloud runs — copilot-cli-release-app[bot] · 2026-10-03
- PhD-turned-founder: LLMs flipped which criteria kill programming tool startups — jimmykoppel · 2026-10-03
- Qwen Code v0.24.7 nightly adds local workspace-agent collaboration — qwen-code-ci-bot · 2026-10-03
- Veteran engineer: AI coding now beats humans on quality, not just speed — facontidavide · 2026-10-03