Best-scoring harness combo burns nearly 1.5x more tokens than Qwen Code
yb2698 · x · 2026-10-07
Continuing the harness-impact experiments, the author changed the simplest variable: the system prompt.
- v1 appends an extra instruction to the existing Pi prompt; v2 replaces it with one minimal sentence
- The best score overall came from the 27B model with Pi — but it consumed nearly 1.5x more tokens than Qwen Code
- Key takeaway: there is no best harness, only a best model-harness pairing, with huge implications for end-user cost and efficiency
Related event: Agent Harnesses Share Nearly Identical Tool Definitions, Analysis Finds(6 posts)→
More from coding & agent
- AI Catches Its Own UI Side-Effect via Screenshot Verification Before Opening PR — DanielLockyer · 2026-10-07
- What's Everyone Using as a Claude Code Multiplexer? Conductor Shines but Lacks Remote Control — genmon · 2026-10-07
- New Tool Launches Local Semantic Search Across Agent Sessions — No Vector DB Needed — ycombinator · 2026-10-07
- Claim: Grokbot Is Now the Best for Long-Running Computer Use with Opus 5.5 — ns123abc · 2026-10-07
- Debating LLM Egress Security: Why Every Tool Call Can Run Through Your MITM Proxy — evilsocket · 2026-10-07
- Developer slams Windsurf's dots: heavy restrictions, poor local setup, likely a push to move data to the cloud — sethlazar · 2026-10-07