GPT 6 Luna fails visual reasoning, worst Western model tester has seen in two years
Afinetheorem · x · 2026-09-25
New Blueprint-Bench results from Andon Labs: GPT 6 Sol slightly beats 5.6 Sol, Grok 4.7 slips vs 4.6, and GPT 6 Luna performs poorly. The tester says Luna's visual reasoning broke down badly — it misses image details and is the worst Western model he's benchmarked in two years, while Opus 5.5 was fantastic.
More from Models
- Dhravya Shah Details How typesafeAI's Jev Improves AI Memory and Context Engineering — blaizedsouza · 2026-09-25
- Community Speculates DeepSeek V4 Coming Soon After Holiday Post From Liang Wenfeng — teortaxesTex · 2026-09-25
- NaceAI launches Drex, a sub-6B decision model that tops the public Decision Index at 51.73 — ordax · 2026-09-25
- Model audit showdown: Astra dominates, Opus and Fable close, Grok 4.7 and GPT-6 Sol lag far behind — ivan_bezdomny · 2026-09-25
- Uncensored local model Bonzai 2 27B tops benchmarks, runs on 12GB VRAM — alexcovo_eth · 2026-09-25
- Same prompt, Opus 5.5 one-shot video generation put to a public retest with different tools — drrickio · 2026-09-25