Stanford and MIT paper: the code harness around an LLM can swing benchmark results up to 6x

burkov · x · 2026-09-15

A Stanford + MIT paper on Model Harnesses shows AI performance depends not just on the model but on the surrounding system code — what gets stored, retrieved, shown to the model, and how the workflow runs. With the same underlying LLM, changing the harness can create up to a 6x performance gap on the same benchmark. Burkov notes in the thread that X no longer penalizes posts with links, so URLs need not be broken up.

Related event: Stanford/MIT Paper: Same Model Can Perform 6x Worse With a Different Harness(3 posts)→

Original post →

More from Research

Research channel →