Same Model, Different Harness: Coding-agent results vary by context policy

rohanpaul_ai · x · 2026-08-31

A study compares Yuj coding agent configurations: a control keeping full history vs. a treatment compressing old tool outputs and detecting stalls. On SWE-bench Verified with a 20k token window, Qwen2.5-Coder's mean F2PF rose from 28% to 49%, and complete solutions increased from 43 to 72. Benefits vanished at 262k tokens. The research suggests evaluating the model and harness as a unified solver.

Related event: Harness Optimization Doubles Qwen's SWE-bench Performance(2 posts)→

Original post →

More from coding & agent

coding & agent channel →