Agent harness engineering: change one variable, test five pipeline stages
MaryamMiradi · x · 2026-09-16
Most people improve AI agents by changing everything at once. Maryam Miradi proposes a 20-minute testing harness and a five-step engineering loop — testing parser, retriever, decision architecture, model, and context configuration one variable at a time. Run the same cases, trace the pipeline, and compare quality, cost, latency, grounding, and pass rates against baseline to localize breakage and catch regressions.
Related event: Harness Engineering: making AI agent improvement a measurable discipline(4 posts)→
More from coding & agent
- Polyphonic details: auto-imported skill libraries, tag-a-into-chat agents, early Mnemos memory system — RileyRalmuto · 2026-09-16
- Polyphonic beta: a Mac 'home' where Claude Code, Codex and other agents share memory and collaborate — RileyRalmuto · 2026-09-16
- Together AI CPO cuts PRDs to 2 pages plus a prototype to fix "context window explosions" — aakashgupta · 2026-09-16
- Agent benchmark turns real incidents into tasks: cached token acted as wrong user for 40 minutes — Kind-Atmosphere9655 · 2026-09-16
- Notification tragedy of the commons: a personal agent should suppress ~95% of them — signulll · 2026-09-16
- skillmem: open-source procedural memory for MCP agents, with hard-won recursion lessons — mrPetrukovich · 2026-09-16