Context overflow fails 78.7% of coding agent tasks; memory fixes lift solve rate to 58%

alex_verem · x · 2026-09-25

A detailed walkthrough of the harness-design study: with a 32k-token window and no memory management, 78.7% of GitHub tasks failed from context overflow (Nemotron-3 550B solved just 6.4%); trimming tool outputs and summarizing earlier steps lifted the same model to 51–58%.

Planning's effect depends on model strength: the 30B model stalled after a median of 5 turns and dropped from 25.2% to 13.6% without a plan, while stronger models kept accuracy within 2 points and cut cost 30%. Tools split similarly—550B did better with a bare terminal at 53% lower cost, while Mistral Medium 3.5 fell from 68.6% to 45.4% without predefined tools.

There is no single best setup; harness components should be chosen per model, task, and budget. Benchmark scores measure model plus harness, and the harness controls much of the number.

Related event: 176 Controlled Experiments Quantify How Agent Harness Design Affects Performance(2 posts)→

Original post →

More from coding & agent

coding & agent channel →