GoBench: GPT-6 Astra tops new Go-playing reasoning benchmark at 2568 Elo
Roland31415 · reddit · 2026-09-17
A Reddit user built GoBench, a benchmark measuring LLM general reasoning via the game of Go. GPT-6 Astra max scores 2568 Elo vs 2076 for Opus 5 high and 1929 for Sol max; with internet-free coding enabled, Codex+Astra reaches 3563 Elo vs 2656 for Codex+Sol. Leaderboard is public.
Related event: GoBench: Testing LLM Reasoning Through the Game of Go(2 posts)→
More from Models
- Humansand Launches Persimmon, a Large-Scale Model Simulating Human Interaction — niloofar_mire · 2026-09-17
- Persimmon launches as first large-scale model to realistically simulate human interaction — niloofar_mire · 2026-09-17
- Dev shares bill: accidental prompt sent to Opus still cost just 30 cents — haydendevs · 2026-09-17
- Code Arena: GPT-6 Astra ranks #1, but Claude Fable 5.1 wins more head-to-head battles — arena · 2026-09-17
- Forcing Jev, a typed decision model that can't output freeform text, to generate anyway — yoimnotkesku · 2026-09-17
- Qwen2.5-1B RLCD fine-tune with parallel constrained decoding trends on Hugging Face — harshatheg · 2026-09-17