GPT-5.6 Three-Tier Benchmark Comparison
docdavkitty · reddit · 2026-07-13
This post provides a comprehensive horizontal evaluation of the three tiers of OpenAI GPT-5.6: Sol / Terra / Luna.
Key Information
- OpenAI released GPT-5.6 GA on July 9
- The three tiers can be iterated independently:
- Sol: $5 / $30 per 1M tokens
- Terra: $2.50 / $15
- Luna: $1 / $6
- Added max reasoning and ultra multi-agent modes
Major Benchmark Performance
- Terminal-Bench 2.1: Sol 88.8%, Terra 87.4%, Luna 84.7%
- BrowseComp: Sol 92.2%, described by the author as SOTA
- AA Coding Agent Index: Sol 80, Terra 77.4, Luna 74.6
- SWE-Bench Pro: Sol 64.6%, though the author notes OpenAI has questioned this benchmark
- DeepSWE value: Luna's cost-effectiveness is considered extremely high, at roughly 24 pts / $1, clearly outperforming more expensive options
Routing Conclusions
- Terra is the default choice for most workloads
- Sol is primarily suited for the hardest agentic / terminal scenarios
- Luna is better for high-throughput pipelines, offering strong cost-effectiveness
- Ultra mode costs about 3x more but only adds 3 points, making it generally not worth the price
More from Models
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11