GPT-6 Luna shows vision regression vs GPT-5.6: extraction drops 81.79% to 66.67%
ducha_aiki · x · 2026-09-24
Developer skalskip92 shared benchmark comparisons claiming "GPT-6 Luna" is a worse vision model than "GPT-5.6 Luna". On high difficulty: extraction fell from 81.79% to 66.67%, counting from 70.72% to 64.41%, and reasoning from 65.56% to 60.71%. Only detection improved, rising from 62.29% to 64.12%. The thread includes multiple failure examples and links to the full evaluation.
More from Models
- ChatGPT Voice gains plugins, ChatGPT Work access, and GPT-6 Astra/Sol/Luna support — romainhuet · 2026-09-24
- GPT-6 Luna beats GPT-5.6 Sol on HealthBench Professional at ~34x lower cost — BorisMPower · 2026-09-24
- Theory: model 'nerfing' may come from mixed heterogeneous inference hardware, not intent — michellechen · 2026-09-24
- "Please Don't Start This with LLMs": Backlash Against Max Prime-Style Model Naming — scaling01 · 2026-09-24
- Limite 1B 'Violetto': tiny open-source model claims competition-math wins over far larger systems — tensorqt · 2026-09-24
- MentalHealthBench: Frontier Models Improving but Gaps Remain in Context-Seeking — thekaransinghal · 2026-09-24