LLM leaderboards are now often measuring the harness too, Gary Marcus warns
GaryMarcus · x · 2026-07-22
Gary Marcus amplifies a point that matters for model evaluation: teams have quietly shifted from benchmarking LLMs alone to benchmarking LLM + harness, while keeping the same leaderboard labels.
The core complaint is methodological: if the system under test changes but the benchmark name does not, comparisons become misleading. Marcus says this is important and people need to understand it to grasp what is happening in the field.
More from Research
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11