GPT-5.6 Three-Tier Benchmark Comparison
docdavkitty · reddit · 2026-07-13
This post provides a comprehensive horizontal evaluation of the three tiers of OpenAI GPT-5.6: Sol / Terra / Luna.
Key Information
- OpenAI released GPT-5.6 GA on July 9
- The three tiers can be iterated independently:
- Sol: $5 / $30 per 1M tokens
- Terra: $2.50 / $15
- Luna: $1 / $6
- Added max reasoning and ultra multi-agent modes
Major Benchmark Performance
- Terminal-Bench 2.1: Sol 88.8%, Terra 87.4%, Luna 84.7%
- BrowseComp: Sol 92.2%, described by the author as SOTA
- AA Coding Agent Index: Sol 80, Terra 77.4, Luna 74.6
- SWE-Bench Pro: Sol 64.6%, though the author notes OpenAI has questioned this benchmark
- DeepSWE value: Luna's cost-effectiveness is considered extremely high, at roughly 24 pts / $1, clearly outperforming more expensive options
Routing Conclusions
- Terra is the default choice for most workloads
- Sol is primarily suited for the hardest agentic / terminal scenarios
- Luna is better for high-throughput pipelines, offering strong cost-effectiveness
- Ultra mode costs about 3x more but only adds 3 points, making it generally not worth the price
More from Models
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22
- Safety Risks of Long-Running Models: OpenAI Shares Codex Alignment Insights — burny_tech · 2026-07-22
- Google launches three new Gemini models, including a cybersecurity system — Polymarket · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22