Stanford/MIT Paper: Same Model Can Perform 6x Worse With a Different Harness
A Stanford/MIT paper introduces 'Model Harnesses': surrounding system code—not just the model—drives performance, with the same model varying up to 6x across harnesses. Their Meta-Harness automatically optimizes harness code, gaining 4.7 points on IMO-level math.
2026-09-13 ~ 2026-09-15 · 3 related posts
- Stanford/MIT paper: same LLM shows up to 6x gap depending on the harness — rohanpaul_ai · 2026-09-13
- Meta-Harness auto-searches harness code, gaining 4.7 points on IMO-level math — zainhas · 2026-09-14
- Stanford and MIT paper: the code harness around an LLM can swing benchmark results up to 6x — burkov · 2026-09-15