Gemini 3.6 Flash matches 3.5 Flash on the AA index and cuts DeepSWE costs by 52%
koltregaskes · x · 2026-07-22
Gemini 3.6 Flash posts mixed but useful benchmark results across coding and general intelligence leaderboards.
- On DeepSWE (long-horizon software engineering) it scores 49%, up from 37% on 3.5 Flash, while average cost drops to $3.53/task from $7.34 and output tokens fall to 97k from 276k.
- In the same DeepSWE band, Claude Opus 4.8 medium also scores 49% at $3.44, GPT-5.6 Sol low scores 45% at $1.07, and Claude Sonnet 5 high scores 48% at $7.43.
- On the Artificial Analysis Intelligence Index, 3.6 Flash scores 50, matching 3.5 Flash but at 280 tokens/sec and with cheaper pricing ($1.50 in / $7.50 out, 17% cheaper than 3.5 Flash).
- The chart also shows 3.6 Flash sitting below top entries like GPT-5.6 Luna max (51), GLM-5.2 max (51), and Claude Sonnet 5 max (53).
Related event: Gemini 3.6 Flash Benchmarks: Faster, Cheaper, Mixed Real-World Tests(5 posts)→
More from Models
- Qwen3.8-Max takes No. 1 on FlashInfer after 500 runs and 30K tool calls — YouJiacheng · 2026-07-22
- Xiaohongshu’s dots-note-3.0 reportedly scores 42/42 on the IMO test — xiaohu · 2026-07-22
- As Frontier Labs Chase AGI, Specialized AI Thrives on Cost and Speed — bendee983 · 2026-07-22
- Grok 4.5 gets praise for clearer, more direct technical writing — elonmusk · 2026-07-22
- OpenAI, Anthropic and Google fall to 83.29% of model spend as Moonshot surges 8.4x — teortaxesTex · 2026-07-22
- Frontier LLMs all lost money in a 1.6-year synthetic trading test — Scobleizer · 2026-07-22