GoBench: LLMs hit 2500 Elo on 9x9 Go vs KataGo's 4400, r=0.83 with ARC-AGI 2
Roland31415 · reddit · 2026-09-17
GoBench evaluates LLMs on 9x9 Go against a ladder of KataGo opponents from random to superhuman:
- Measures general reasoning ability, strongly correlated with ARC-AGI 2 (r=0.83), and far from saturated
- GPT-6 Astra max reaches 2500 Elo, well below top KataGo's 4400
- With coding tools and two hours of prep, Codex with Astra reaches 3560 Elo, showing tool use narrows the gap
Leaderboard, code and paper are public; the author pledges to keep updating the board.
Related event: GoBench: Testing LLM Reasoning Through the Game of Go(2 posts)→
More from Models
- User burns 4 Codex banked resets in 30 minutes to reset 5-hour limits back-to-back — flowersslop · 2026-09-17
- Stealth model Union Alpha scores 74% on DeepSWE, beating GPT-5.6 Sol at lower cost — ZeroStateReflex · 2026-09-17
- $200 Pro user keeps hitting image rate limit popup that doesn't actually block — Jello_Hello_Fellos · 2026-09-17
- GoBench: GPT-6 Astra tops new Go-playing reasoning benchmark at 2568 Elo — Roland31415 · 2026-09-17
- GPT-6 Astra lands first Slay the Spire 2 A10 win on stream with a Demon Form deck — Jsevillamol · 2026-09-17
- Bio speaker still mocks ChatGPT hallucinations; author asks if they even used deep research — zebird0 · 2026-09-17