Gary Marcus uses ChessBench to argue LLMs are still weak at chess

GaryMarcus · x · 2026-07-29

Gary Marcus mocked the idea that current LLMs are anywhere near “smarter than the smartest humans,” pointing to chess as a simple counterexample.

He cited @Kasparov63’s peak rating of 2851 and argued that Kasparov might have beaten the best commercial LLMs in chess even as a 7-year-old. The post also references ChessBench, a new benchmark for language models, which visualizes how different models score on coherence, accuracy, and Elo in chess-related tasks.

In other words: the benchmark is being used to underline a broader point that language models still fall far short of robust game-playing competence, despite AGI hype.

Original post →

More from Models

Models channel →