Gemini 3.6 Flash matches 3.5 Flash on the AA index and cuts DeepSWE costs by 52%
koltregaskes · x · 2026-07-22
Gemini 3.6 Flash posts mixed but useful benchmark results across coding and general intelligence leaderboards.
- On DeepSWE (long-horizon software engineering) it scores 49%, up from 37% on 3.5 Flash, while average cost drops to $3.53/task from $7.34 and output tokens fall to 97k from 276k.
- In the same DeepSWE band, Claude Opus 4.8 medium also scores 49% at $3.44, GPT-5.6 Sol low scores 45% at $1.07, and Claude Sonnet 5 high scores 48% at $7.43.
- On the Artificial Analysis Intelligence Index, 3.6 Flash scores 50, matching 3.5 Flash but at 280 tokens/sec and with cheaper pricing ($1.50 in / $7.50 out, 17% cheaper than 3.5 Flash).
- The chart also shows 3.6 Flash sitting below top entries like GPT-5.6 Luna max (51), GLM-5.2 max (51), and Claude Sonnet 5 max (53).
Related event: Gemini 3.6 Flash Review: Faster and Cheaper, But Not Smarter(16 posts)→
More from Models
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11