Why do commercial LLMs still fail at chess despite all that pretraining?

alex_peys · x · 2026-07-29

The author argues that chess is still a useful benchmark for LLMs because commercial models continue to perform poorly at it despite having likely seen huge amounts of chess material in pretraining.

Their key questions are:

The post is less a claim than a research-style question about what pretraining actually contains and how capabilities can be unlocked.

Related event: Poor Chess Performance in Commercial LLMs Highlights Capability Bottlenecks(2 posts)→

Original post →

More from Research

Research channel →