120-run eval: coding agents complete 0/40 MCP workflows without instructions
colinmcnamara · x · 2026-10-06
Jeff Linwood presented a controlled experiment at Austin AI MUG using the Harbor framework, showing that instructions are decisive for coding agents' MCP tool use.
Setup:
- 5 task fixtures × 2 agents (Claude Code/Codex) × 2 model tiers × 3 instruction conditions × 2 repetitions = 120 sandboxed runs
- Coding success measured with hidden pytest tests; workflow outcome verified via Zebric audit logs and final database state
Results:
- All 120 runs passed the coding tests
- No instructions: 0/40 runs completed the task workflow (agents never pulled tasks from the board)
- Short instruction (list tasks → mark inprogress → mark done): 20/40 completed
Takeaway: agents don't lack tool ability — they lack knowledge of when to use tools; embedding task-tracker tool descriptions into repo instruction files is key to real workflow execution.
More from coding & agent
- A Planted 'P.S.' Fooled Jev, TypeSafe's New Decision Model — a Simple Rule Caught It — Internal-Lie-5197 · 2026-10-06
- Claude Code's New Dreaded Message: 'Compacted While Idle, Before the Prompt Cache Expired' — dSebastien · 2026-10-06
- Dev uses 5 Grok agents over 3 days to pre-generate a full AI dictionary for English readers — sujingshen · 2026-10-06
- Student builds dAIly, a local-first AI planning agent powered by Gemma — damnGruz · 2026-10-06
- "It's all prompts": a harness is just a tool for restructuring input and output tokens — shakoistsLog · 2026-10-06
- Built in Three Prompts and an Hour: Why "I Wanted This to Exist" Now Justifies Software — TheMoonMidas · 2026-10-06