LLM leaderboards are now often measuring the harness too, Gary Marcus warns
GaryMarcus · x · 2026-07-22
Gary Marcus amplifies a point that matters for model evaluation: teams have quietly shifted from benchmarking LLMs alone to benchmarking LLM + harness, while keeping the same leaderboard labels.
The core complaint is methodological: if the system under test changes but the benchmark name does not, comparisons become misleading. Marcus says this is important and people need to understand it to grasp what is happening in the field.
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22