GPT-6 Luna shows vision regression vs GPT-5.6: extraction drops 81.79% to 66.67%

ducha_aiki · x · 2026-09-24

Developer skalskip92 shared benchmark comparisons claiming "GPT-6 Luna" is a worse vision model than "GPT-5.6 Luna". On high difficulty: extraction fell from 81.79% to 66.67%, counting from 70.72% to 64.41%, and reasoning from 65.56% to 60.71%. Only detection improved, rising from 62.29% to 64.12%. The thread includes multiple failure examples and links to the full evaluation.

Original post →

More from Models

Models channel →