Context overflow fails 78.7% of coding agent tasks; memory fixes lift solve rate to 58%
alex_verem · x · 2026-09-25
A detailed walkthrough of the harness-design study: with a 32k-token window and no memory management, 78.7% of GitHub tasks failed from context overflow (Nemotron-3 550B solved just 6.4%); trimming tool outputs and summarizing earlier steps lifted the same model to 51–58%.
Planning's effect depends on model strength: the 30B model stalled after a median of 5 turns and dropped from 25.2% to 13.6% without a plan, while stronger models kept accuracy within 2 points and cut cost 30%. Tools split similarly—550B did better with a bare terminal at 53% lower cost, while Mistral Medium 3.5 fell from 68.6% to 45.4% without predefined tools.
There is no single best setup; harness components should be chosen per model, task, and budget. Benchmark scores measure model plus harness, and the harness controls much of the number.
More from coding & agent
- AMD's software VP on ROCm's open toolchain and AI agents writing GPU code — AnushElangovan · 2026-09-25
- Building agents for cloud infra: why one bad terraform destroy beats a bad email — Helpful-Man64 · 2026-09-25
- Everyone solved AI code review, nobody solved what happens after merge — trvklhn666 · 2026-09-25
- When AI agents check out off-store, what proof do merchants need to trust it? — ConvertMyStore · 2026-09-25
- Stop Fixing AI Mistakes: Feed Models Your Context, Says Dev Strategy Thread — chaseleantj · 2026-09-25
- fable-advisor ships setup command to max out Claude and ChatGPT subs at once — daniel_mac8 · 2026-09-25