Same Model, Different Harness: Coding-agent results vary by context policy
rohanpaul_ai · x · 2026-08-31
A study compares Yuj coding agent configurations: a control keeping full history vs. a treatment compressing old tool outputs and detecting stalls. On SWE-bench Verified with a 20k token window, Qwen2.5-Coder's mean F2PF rose from 28% to 49%, and complete solutions increased from 43 to 72. Benefits vanished at 262k tokens. The research suggests evaluating the model and harness as a unified solver.
Related event: Harness Optimization Doubles Qwen's SWE-bench Performance(2 posts)→
More from coding & agent
- TablePro: Open Source Database Client with MCP Support and AI Chat — tom_doerr · 2026-08-31
- Using AI Clairvoyance for game AI opponent evaluation — draginol · 2026-08-31
- Manzanas: Control 7 iOS Sims Across 3 MacBooks in Real Time for Agents — Plastic-Risk-6309 · 2026-08-31
- Engineers Share Scars From Massive AI Production Bill Spikes — BasePsychological899 · 2026-08-31
- How to handle parallel AI coding sessions in the same repo? — McButterblump · 2026-08-31
- Using Google Drive as ChatGPT's external operational memory — edalgomezn · 2026-08-31