21 model-harness pairs tested: framework choice barely moves success rate but swings cost
iScienceLuvr · x · 2026-09-19
A new analysis compares 21 model-harness pairs across 7 models and 3 harnesses and lands on three findings:
- Harness choice has little effect on task success rate, but can significantly change cost
- A simple harness can be competitive with more complex setups
- Models sometimes perform better on other harnesses than their own
The author makes a fairly compelling case that a simpler harness like Pi may be the better choice — for anyone picking an AI coding/agent framework, the takeaway is to prioritize simplicity and cost over feature bloat.
More from coding & agent
- 8 components of harness engineering: why the system around the model, not the model, makes agents reliable — blaizedsouza · 2026-09-19
- Google's 12-page Agentic Engineering guide lays out a 5-stage pipeline for building agent teams — blaizedsouza · 2026-09-19
- AgentRun: a purpose-built harness for high-frequency repetitive knowledge work — blaizedsouza · 2026-09-19
- Qwen 3.8 27B on a single RTX 5090 builds a full animation using only code — Acceptable-Object390 · 2026-09-19
- Compiler Explorer runs 92M compilations a year — here's how it works — blaizedsouza · 2026-09-19
- Netflix builds module-first, agent-friendly Java tooling on OpenJDK foundations — blaizedsouza · 2026-09-19