GoBench: Testing LLM Reasoning Through the Game of Go
Roland31415 released GoBench, which tests LLM reasoning by playing 9x9 Go against KataGo ladders, correlating 0.83 with ARC-AGI 2, with GPT-6 Astra leading at 2568 Elo.
2026-09-17 ~ 2026-09-17 · 2 related posts
- GoBench: LLMs hit 2500 Elo on 9x9 Go vs KataGo's 4400, r=0.83 with ARC-AGI 2 — Roland31415 · 2026-09-17
- GoBench: GPT-6 Astra tops new Go-playing reasoning benchmark at 2568 Elo — Roland31415 · 2026-09-17