Stanford and MIT paper: the code harness around an LLM can swing benchmark results up to 6x
burkov · x · 2026-09-15
A Stanford + MIT paper on Model Harnesses shows AI performance depends not just on the model but on the surrounding system code — what gets stored, retrieved, shown to the model, and how the workflow runs. With the same underlying LLM, changing the harness can create up to a 6x performance gap on the same benchmark. Burkov notes in the thread that X no longer penalizes posts with links, so URLs need not be broken up.
More from Research
- Jay Alammar at PyData: ten questions that explain benchmark score gaps — JayAlammar · 2026-09-15
- RLE-Bench: 48 Tasks Benchmark LLM Agents on Full Robotics Engineering, Not Just Control — daibond_alpha · 2026-09-15
- BrainGPT creator: AI discoveries will soon be as explainable as quantum mechanics to a dog — every · 2026-09-15
- Ethan Mollick tests GPT on Linear A, warns it's 'researchslop' until GPT-6 cracks it — emollick · 2026-09-15
- Benchmark tests 71 AI personal assistants across 15 dimensions, Muse tops at 9.1 — ai · 2026-09-15
- FlyWire publishes DIY guide to simulate a fly brain with ~160k neurons — patrickmineault · 2026-09-15