GPT-6 Sol/Luna benchmarks: cheaper and less hallucination, but weaker than GPT-5.6 on some tasks

After the release of GPT-6 Sol and Luna, Artificial Analysis added both models to its Intelligence Index v4.3 comparison and published a series of in-depth data. Meanwhile, multiple users reported in hands-on testing that GPT-6 Sol underperforms the previous-generation GPT-5.6 Sol on some complex tasks. The overall picture: the new models are cheaper, more efficient, and hallucinate less, but they do not comprehensively surpass their predecessors, with regressions on software engineering and knowledge-work benchmarks.

Confirmed

Unconfirmed

Why it matters

It is unusual for a new model not to beat its predecessor across all metrics, prompting discussion about whether generational progress is slowing and whether vendors selectively disclose benchmarks. At the same time, halved prices combined with sharply lower hallucination rates remain highly attractive for cost-sensitive production use. Directional divergence between users' side-by-side tests and third-party benchmarks suggests that testing against your own workload matters more than blindly trusting "the new generation."

2026-09-23 ~ 2026-09-23 · 12 related posts

Primary sources