GPT-6 Sol and Luna Score Below GPT-5.6: First Major Release Weaker Than Its Predecessor

srchvrs · x · 2026-09-24

On the Extended NYT Connections benchmark (940 puzzles with added trick words), Opus 5.5, GPT-6 Sol, and GPT-6 Luna all score slightly worse but far cheaper than their predecessors. Commenters note this may be the first major model release in modern history that is less capable than the previous generation, reversing a trend since GPT-3.

Original post →

More from Models

Models channel →