GPT-6 without code execution beats GPT-5.6 with Python on MazeBench
patience_cave · x · 2026-09-07
Wrap-up of the MazeBench thread: GPT-6 Astra scored 14% without code execution, beating GPT-5.6 Sol's 13% even when the latter had Python available — a clear generational leap, yet still far from the eval's hardest challenges. The MazeBench leaderboard is public at mazebench.com.
More from Models
- Google's Astra Agent Allegedly Crushes Existing CAPTCHAs, Sparking Rethink of Bot Checks — eyishazyer · 2026-09-07
- Intelligence price collapsed ~1000x in 18 months as models keep getting smarter — ccerrato147 · 2026-09-07
- Blogger says Google Astra's natural tone broke his last dependency on Claude — StewartalsopIII · 2026-09-07
- VC slams frontier lab subscriptions: usage limits to worsen, top models to shrink — StewartalsopIII · 2026-09-07
- GPT Astra scores just 13% on MazeBench without tools — Wonderful_Buffalo_32 · 2026-09-07
- GPT-6 Astra Generates Stunning Math Animation in One Shot — omarsar0 · 2026-09-07