176-Setting Study Dissects Coding Agent Harnesses: Context Management Matters Most on Tight Budgets
_akhaliq · x · 2026-09-19
A new paper, "An Empirical Study of Harness Design for Coding Agents," builds a modular harness that keeps the execution loop fixed while varying planning, tool interfaces, and context management, evaluated across 176 settings with four models on SWE-Bench Verified and Terminal-Bench 2.1.
Key findings:
- Context management matters most with tight context budgets: it prevents overflow from killing runs early; benefits shrink as windows grow.
- Staging elision before LLM summarization is the most efficient strategy, controlling peak context with similar success rates.
- Planning shifts from accuracy scaffold to efficiency aid: for weak models it keeps trajectories alive; for strong models it mainly cuts redundant post-edit verification.
- Predefined tools help bash-weak models; bash-only interfaces cut cost for bash-capable ones.
Related event: 176 Experiments Show Context Management Is Key for Coding Agents(3 posts)→
More from coding & agent
- mitsuhiko: team gave web components an honest chance, went all in on React — mitsuhiko · 2026-09-19
- jev runs ~14.8x faster than DeepSeek+vision on real-device Android e2e testing — kevinkern · 2026-09-19
- -render + jev experiment shows instant Generative UI rendered in milliseconds — cramforce · 2026-09-19
- Stop Building AI Agents Like Coding Agents: A 6-Step Framework for Knowledge Work — MaryamMiradi · 2026-09-19
- Practical Jev uses: classifying prompts by tier and auto-detecting refusals — QuixiAI · 2026-09-19
- Commons launches: open workspaces where humans and AI agents run autonomous organizations — Scobleizer · 2026-09-19