Gemini 3.6 Flash beats 3.5 Flash, but lags GPT-5.6 Sol and Terra on cost and quality
haider1 · x · 2026-07-23
The post says Gemini 3.6 Flash is a confusing release: it is a clear upgrade over 3.5 Flash, but still trails GPT-5.6 Sol and Terra on intelligence while costing more and using more tokens.
The attached chart compares DeepSWE Pass@1, average cost per task, and output tokens across effort levels. In the high-effort and xHigh settings, GPT-5.6 Sol and Terra lead on quality-efficiency tradeoffs, while Gemini 3.6 Flash lands below them on both performance and cost efficiency.
More from Models
- Anthropic says a cutoff-date bug showed March 2026 in some domains — _arohan_ · 2026-07-23
- A simple quant benchmark could expose frontier-model failures fast — PtrPomorski · 2026-07-23
- Anthropic Adjusts Subscription Tiers: PRO Loses Fable 5 Access After Credits — shaunralston · 2026-07-23
- Inkling scores 836 Elo on AA-Briefcase, trailing top open-weight models — ArtificialAnlys · 2026-07-23
- OpenAI Model Hacks HuggingFace Using Zero-Day Exploit During Benchmark — Gary Marcus · 2026-07-23
- Security Differences Between Closed and Open Source Models: Insights from OpenAI's Escape Incident — robleclerc · 2026-07-23