GPT-6 Sol/Luna benchmarks: cheaper and less hallucination, but weaker than GPT-5.6 on some tasks
After the release of GPT-6 Sol and Luna, Artificial Analysis added both models to its Intelligence Index v4.3 comparison and published a series of in-depth data. Meanwhile, multiple users reported in hands-on testing that GPT-6 Sol underperforms the previous-generation GPT-5.6 Sol on some complex tasks. The overall picture: the new models are cheaper, more efficient, and hallucinate less, but they do not comprehensively surpass their predecessors, with regressions on software engineering and knowledge-work benchmarks.
Confirmed
- Pricing halved: Artificial Analysis data shows Sol dropping from $4/$20 to $2/$10 per million tokens, and Luna dropping from $0.20/$1.20 to roughly half.
- Coding agent gains: In the Codex environment, GPT-6 Sol (max) scored 57 on the Coding Agent Index, up 2 points over GPT-5.6, at about half the per-task cost of its predecessor.
- Lower hallucination rates: On the AA-Omniscience benchmark, Sol's hallucination rate fell from 92% to 60%, and Luna's from 93% to 77% (max effort).
- Knowledge-work regression: On GDPval-AA v2.1 (adapted from an OpenAI dataset covering 44 occupations) in max effort mode, both Sol and Luna regressed; manual review found the main cause was shorter deliverables that missed key points.
- DeepSWE decline: Hands-on tests and reports from multiple users (power97992, Angaisb, etc.) show GPT-6 Sol scoring lower than GPT-5.6 Sol on DeepSWE.
Unconfirmed
- A claim relayed by user ns123abc says GPT-6 Sol is also worse than 5.6 Sol at computer use and research debugging, with no improvement on cybersecurity tasks; the poster themselves said it is unverified.
- weswinder and adonissingh suspect the vendor deliberately avoided and "buried" the DeepSWE benchmark; this is speculation with no official response to back it.
- User power97992 reported that GPT-6 Sol's xhigh tier loses to 5.6 Sol's max tier on some coding tasks—personal testing with a limited sample.
Why it matters
It is unusual for a new model not to beat its predecessor across all metrics, prompting discussion about whether generational progress is slowing and whether vendors selectively disclose benchmarks. At the same time, halved prices combined with sharply lower hallucination rates remain highly attractive for cost-sensitive production use. Directional divergence between users' side-by-side tests and third-party benchmarks suggests that testing against your own workload matters more than blindly trusting "the new generation."
2026-09-23 ~ 2026-09-23 · 12 related posts
Primary sources
- GPT-6 Sol scores slightly below GPT-5.6 Sol on DeepSWE, only cheaper — Angaisb_ · 2026-09-23
- [source] Every's GPT-6 Sol Hands-On: Near-Astra Writing, 50% Cheaper Than 5.6 Sol — every · 2026-09-23
- [source] GPT-6 Sol and Luna Halve Prices, Sol Cuts Hallucination Rate from 92% to 60% — ArtificialAnlys · 2026-09-23
- GPT-6 Sol Gains 2 Points in Coding Agent Index at Half the Cost — ArtificialAnlys · 2026-09-23
- GPT-6 Models Hallucinate Substantially Less Than Predecessors — ArtificialAnlys · 2026-09-23
- GPT-6 Regresses on GDPval Knowledge-Work Benchmark Due to Shorter Deliverables — ArtificialAnlys · 2026-09-23
- GPT-6 Sol and Luna added to Artificial Analysis Intelligence Index v4.3 — ArtificialAnlys · 2026-09-23
- Artificial Analysis Breaks Down Individual Evals in Intelligence Index v4.3 — ArtificialAnlys · 2026-09-23
- GPT-6 Sol Reportedly Worse Than 5.6 Sol on DeepSWE and Debugging — ns123abc · 2026-09-23
- [source] Hands-on: GPT-6 Sol xhigh feels worse than 5.6 Sol Max at coding — power97992 · 2026-09-23
- GPT-6 sol scores worse than GPT-5.6 sol on DeepSWE, sparking benchmark trust concerns — adonis_singh · 2026-09-23
- GPT-6 Sol Weaker Than 5.6 Sol on Complex Tasks, But Wins on Cost and Efficiency — FalconsArentReal · 2026-09-23