GPT-6 Sol/Luna benchmarks leak: near-Fable performance at 80-85% lower cost
kimmonismus · x · 2026-09-23
kimmonismus shares GPT-6 Sol & Luna benchmark numbers:
- FrontierCode: Sol 48.4% vs Fable 5.1's 48.7% (both xhigh), $1.37 vs $9.27 per task (85% cheaper)
- DeepSWE: 68.8% at max vs Fable 5's best 69.9% at xhigh, $2.74 vs $13.41 (80% cheaper)
- AutomationBench: 33.2% at xhigh beats Fable 5.1 + Opus 5 fallback's 31.4%, $0.27 vs $2.45+ per task
Bottom line: GPT-6 series matches near-frontier performance at dramatically lower cost.
Related event: OpenAI Launches GPT-6 Sol and Luna at Half the Price(48 posts)→
More from Models
- "People are loving 5.5": Matt Shumer congratulates Anthropic on new model reception — mattshumer_ · 2026-09-23
- 10 Claude agents spend 15 hours devising and Lean-proving a faster shortest-path algorithm, C-HD — ctjlewis · 2026-09-23
- Developer claims Anthropic noticed community posts about 'Claudelish' speech quirks — evijit · 2026-09-23
- scaling01 taunts haters after joking Anthropic won't ship Opus 5.5 and Sol today — scaling01 · 2026-09-23
- Delip Rao downgrades from $200/mo Google One Ultra to $50 Pro, leaning on local models — deliprao · 2026-09-23
- Why AI progress accelerated: Claude 4.5 kicked off narrow RSI and open-weight catch-up — maksym_andr · 2026-09-23