Gary Marcus uses ChessBench to argue LLMs are still weak at chess
GaryMarcus · x · 2026-07-29
Gary Marcus mocked the idea that current LLMs are anywhere near “smarter than the smartest humans,” pointing to chess as a simple counterexample.
He cited @Kasparov63’s peak rating of 2851 and argued that Kasparov might have beaten the best commercial LLMs in chess even as a 7-year-old. The post also references ChessBench, a new benchmark for language models, which visualizes how different models score on coherence, accuracy, and Elo in chess-related tasks.
In other words: the benchmark is being used to underline a broader point that language models still fall far short of robust game-playing competence, despite AGI hype.
More from Models
- Low API Price ≠ Cheap: Kimi K3's Low Cache Hit Rate Hurts Real-World Costs — peterjliu · 2026-07-29
- Kimi K3 Takes #1 in Code Arena Fullstack, Beating GPT-5.6 and Claude — KickLassChewGum · 2026-07-29
- Kimi K3 is the First Open-Source Model to Pass Compound's Internal Benchmark — peterjliu · 2026-07-29
- Claude Opus 5 tops LisanBench while using far fewer tokens in medium mode — scaling01 · 2026-07-29
- Bindu Reddy says OpenAI and Anthropic are fear-mongering over Kimi K3 — bindureddy · 2026-07-29
- Claude CoT leak joke turns Anthropic’s “openness” into a model-bashing meme — OwariDa · 2026-07-29