GPT-6 Sol and Luna Score Below GPT-5.6: First Major Release Weaker Than Its Predecessor
srchvrs · x · 2026-09-24
On the Extended NYT Connections benchmark (940 puzzles with added trick words), Opus 5.5, GPT-6 Sol, and GPT-6 Luna all score slightly worse but far cheaper than their predecessors. Commenters note this may be the first major model release in modern history that is less capable than the previous generation, reversing a trend since GPT-3.
More from Models
- Leak Chatter: Frontier Model 8 Months Ahead of Astra, Only 2 Months Ahead of Trend — scaling01 · 2026-09-24
- Apple Open-Sources LensVLM-9B, a Qwen3.5-9B Finetune That Shrinks Long Docs into Page Images — victormustar · 2026-09-24
- OpenAI power users burn billions of tokens daily; Navier-Stokes project hits trillions of input tokens — scaling01 · 2026-09-24
- Navier-Stokes proof burned 130B output and trillions of input tokens, dwarfing top OpenAI users' few billion daily — scaling01 · 2026-09-24
- Meta Muse scores 4.88 across 40k ratings — analyst says product, distribution, price make it near-unbeatable — RihardJarc · 2026-09-24
- Audi 3D Modeling Benchmark Rerun: Only GPT-6 Astra and Opus 5.5 Are Worth Taking Seriously — kevinkern · 2026-09-24