Open experiments: swapping the system prompt swings coding agent scores by 10%

yb2698 · x · 2026-10-09

yb2698 shares open experiments on coding agent harnesses: on the ALE near-term subset, minimalist harness Pi beats Qwen-code on the same 27B model while consuming 1.5x more tokens. Merely changing the system prompt or tool definitions can swing results by 10% and lengthen successful trajectories. He poses open questions: what makes a good harness, can harnesses be stochastic, do 60+ tool harnesses suit small models, and will minimalist harnesses win long term — calling for more open experiments.

Original post →

More from coding & agent

coding & agent channel →