GPT 6 Luna fails visual reasoning, worst Western model tester has seen in two years

Afinetheorem · x · 2026-09-25

New Blueprint-Bench results from Andon Labs: GPT 6 Sol slightly beats 5.6 Sol, Grok 4.7 slips vs 4.6, and GPT 6 Luna performs poorly. The tester says Luna's visual reasoning broke down badly — it misses image details and is the worst Western model he's benchmarked in two years, while Opus 5.5 was fantastic.

Original post →

More from Models

Models channel →