Artificial Analysis Reviews GPT-6: Halved Prices, Lower Hallucination Rates, but GDPval Regression
On 09-23, Artificial Analysis released an in-depth evaluation of GPT-6 Sol and Luna. Both models have been added to its Intelligence Index v4.3 benchmark suite, where users can compare them item by item against other leading models. The core takeaways: pricing cut across the board, hallucination rates significantly down, better cost-effectiveness for coding agents, but a regression on the GDPval knowledge-work benchmark.
Confirmed
- Halved pricing: Sol dropped from $4/$20 to $2/$10 (per million tokens), and Luna from $0.20/$1.20 to the $0.1 tier
- AA-Omniscience benchmark (max effort) hallucination rates: Sol fell from 92% to 60%, Luna from 93% to 77%, both clearly below the previous generation
- Coding Agent Index in the Codex environment: GPT-6 Sol (max) scored 57, up 2 points from GPT-5.6, at roughly half the per-task cost of its predecessor
- Both models entered Intelligence Index v4.3, and AA also published detailed breakdowns for each sub-benchmark
Unconfirmed
- On GDPval-AA v2.1 (adapted from OpenAI's dataset of economic-value tasks spanning 44 occupations), both Sol and Luna regressed in max effort mode; AA's manual review found the main causes were shorter deliverables and missed key points, but whether this regression affects real-world productivity scenarios still needs further validation
Why it matters
- Halved prices combined with improved hallucination rates mean GPT-6 advances simultaneously on cost-effectiveness and reliability, with direct impact for developers sensitive to API costs
- The GDPval regression stands in contrast to the coding and hallucination improvements, suggesting model capability gains are uneven—before choosing, users should consult the per-category data against their own task types
2026-09-23 ~ 2026-09-23 · 6 related posts
Primary sources
- GPT-6 Sol and Luna Halve Prices, Sol Cuts Hallucination Rate from 92% to 60% — ArtificialAnlys ·
- GPT-6 Models Hallucinate Substantially Less Than Predecessors — ArtificialAnlys ·
- GPT-6 Regresses on GDPval Knowledge-Work Benchmark Due to Shorter Deliverables — ArtificialAnlys ·
- [source] GPT-6 Sol and Luna Halve Prices, Sol Cuts Hallucination Rate from 92% to 60% — ArtificialAnlys · 2026-09-23
- GPT-6 Sol Gains 2 Points in Coding Agent Index at Half the Cost — ArtificialAnlys · 2026-09-23
- [source] GPT-6 Models Hallucinate Substantially Less Than Predecessors — ArtificialAnlys · 2026-09-23
- [source] GPT-6 Regresses on GDPval Knowledge-Work Benchmark Due to Shorter Deliverables — ArtificialAnlys · 2026-09-23
- GPT-6 Sol and Luna added to Artificial Analysis Intelligence Index v4.3 — ArtificialAnlys · 2026-09-23
- Artificial Analysis Breaks Down Individual Evals in Intelligence Index v4.3 — ArtificialAnlys · 2026-09-23