Benchmark scorecard pits GPT-5.6 Sol, Claude Fable 5, and Gemini 3.6 Flash
iruletheworldmo · x · 2026-07-22
Frontier scorecard pits GPT-5.6 Sol, Claude Fable 5, and Gemini 3.6 Flash
A benchmark scorecard compares three frontier models across coding, agentic, and computer-use evaluations.
- GPT-5.6 Sol leads on DeepSWE v1.1 and Terminal-Bench 2.1.
- Claude Fable 5 leads on SWE-Bench Pro and OSWorld-Verified.
- Gemini 3.6 Flash is much cheaper than the others, with $1.50 input / $7.50 output per 1M tokens, but trails on several coding and agentic benchmarks.
- The chart also shows tied or near-tied results on some reasoning and tool-use tasks, plus a GDM-MRCR v2 lead for GPT-5.6 Sol.
The image presents raw public benchmark results and pricing side by side, making the tradeoff between cost and capability the main takeaway.
Related event: Gemini 3.6 Flash Benchmarks: Stagnant Intelligence but Improved Efficiency(20 posts)→
More from Models
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22
- Safety Risks of Long-Running Models: OpenAI Shares Codex Alignment Insights — burny_tech · 2026-07-22
- Google launches three new Gemini models, including a cybersecurity system — Polymarket · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22