LLM leaderboards are now often measuring the harness too, Gary Marcus warns
GaryMarcus · x · 2026-07-22
Gary Marcus amplifies a point that matters for model evaluation: teams have quietly shifted from benchmarking LLMs alone to benchmarking LLM + harness, while keeping the same leaderboard labels.
The core complaint is methodological: if the system under test changes but the benchmark name does not, comparisons become misleading. Marcus says this is important and people need to understand it to grasp what is happening in the field.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11