GPT-6 Sol/Luna Benchmarks Leak: ~85% Cheaper Per Task Than Fable 5.1
kimmonismus · x · 2026-09-23
kimmonismus shares benchmark and cost numbers for GPT-6 Sol and Luna:
- FrontierCode: Sol hits 48.4% vs Fable 5.1's 48.7% (both xhigh), at $1.37 vs $9.27 per task — roughly 85% cheaper
- DeepSWE: 68.8% (max) vs Fable 5's best 69.9% (xhigh), $2.74 vs $13.41 per task, 80% cheaper
- AutomationBench: 33.2% (xhigh) vs Fable 5.1 with Opus 5 fallback at 31.4% (max), $0.27 vs at least $2.45 per task
The takeaway: the GPT-6 generation matches near-flagship performance on agentic/coding tasks at a fraction of the per-task cost.
Related event: OpenAI Launches GPT-6 Sol and Luna at Half the Price(48 posts)→
More from Models
- "People are loving 5.5": Matt Shumer congratulates Anthropic on new model reception — mattshumer_ · 2026-09-23
- 10 Claude agents spend 15 hours devising and Lean-proving a faster shortest-path algorithm, C-HD — ctjlewis · 2026-09-23
- Developer claims Anthropic noticed community posts about 'Claudelish' speech quirks — evijit · 2026-09-23
- scaling01 taunts haters after joking Anthropic won't ship Opus 5.5 and Sol today — scaling01 · 2026-09-23
- Delip Rao downgrades from $200/mo Google One Ultra to $50 Pro, leaning on local models — deliprao · 2026-09-23
- Why AI progress accelerated: Claude 4.5 kicked off narrow RSI and open-weight catch-up — maksym_andr · 2026-09-23