GPT-6.1 Sol Scores Just 9% on MazeBench, Barely Beating Opus 5.5
patience_cave · x · 2026-10-02
MazeBench results show GPT-6.1 Sol scoring just 9%, slightly above Opus 5.5. Compared to GPT-6 Astra, Sol retained its puzzle-solving capabilities but at 10% of the cost.
More from Models
- Unverified: Gemini 4 Argon Rumored to Output 1M Tokens in a Single Run — jocarrasqueira · 2026-10-02
- Third-party eval: Cloudflare's clef beats Jev on coding but is too pricey — airesearch12 · 2026-10-02
- GPT-6 Astra Vibe Check: Strong Writing and Design, but Fable Still Builds Better Products — every · 2026-10-02
- MIT Tech Review Argues LLMs Don't Actually Reason – They Pattern-Match — leopoldj · 2026-10-02
- JevBench v1.5.5: top 10 unchanged, Jev still #1 on capability but only 3rd composite — airesearch12 · 2026-10-02
- Anthropic Publishes 'Claude-Shaped Science' Research on AI in Scientific Work — famouswaffles · 2026-10-02