GPT-6.1 Sol Scores Just 9% on MazeBench, Barely Beating Opus 5.5

patience_cave · x · 2026-10-02

MazeBench results show GPT-6.1 Sol scoring just 9%, slightly above Opus 5.5. Compared to GPT-6 Astra, Sol retained its puzzle-solving capabilities but at 10% of the cost.

Original post →

More from Models

Models channel →