Testing 7 models across Claude Code, Codex and Pi: native harness isn't always best
CShorten30 · x · 2026-09-17
An evaluation ran 7 models inside Claude Code, Codex and Pi, with three surprising findings:
- Harness choice has little effect on task success rate but can significantly affect cost.
- A simple harness can be competitive.
- The native harness isn't always the best.
With millions using coding agents, harness selection is an overlooked variable; the authors argue "harness-routing" deserves attention.
More from coding & agent
- Bash is no longer all you need for reliable agent tool calling — yenkel · 2026-09-17
- Dev lead running 50 deploys a day shares how to write and run tests in the vibe coding era — dotey · 2026-09-17
- YC built AI versions of its partners on GLM-5.2, cutting latency 31% vs OpenAI — ycombinator · 2026-09-17
- Shape Claude's Tools the Way You Want Instead of Hiding Behind Indirection — trq212 · 2026-09-17
- Out of Codex quota? Signing in with a new Plus account preserves all your chats — TheMoonMidas · 2026-09-17
- Vercel's fx may switch safety reviewer to Jev: 5-18x faster than GPT Luna — thesaraharminta · 2026-09-17