Gemini 3.8 Flash scores 59 on AA Index, cheapest model at its intelligence at $0.58 per task
ArtificialAnlys · x · 2026-09-02
Artificial Analysis benchmarks Gemini 3.8 Flash, Google's fourth Flash model in under four months:
- Scores 59 on the AA Intelligence Index (high reasoning), up 3 points from 3.7 Flash, matching GPT-5.6 Sol (xhigh) and Grok 4.6 (medium)
- Keeps 3.7 Flash discounted pricing ($0.75/$3.75 per M tokens) through year-end; $0.58 per task, cheapest at its intelligence level but 40% costlier than its predecessor due to 30% more output tokens (48k avg) and more agentic turns
- Gains driven by agentic evals: +12 points on 𝜏³-Banking to 45%
- 300 tok/s, 2.5 min per task (high reasoning); 0.8 min on low reasoning
- 1M context, text/image/video/speech input
More from Models
- 3.8 Flash hits the Pareto frontier: 5x cheaper, 2x faster than Opus 5 at similar intelligence — davidtsong · 2026-09-03
- Pedro Domingos blasts OpenAI: looped transformers aren't more opaque, CoT transparency is fiction — pmddomingos · 2026-09-03
- Google Search Console social properties show bizarre zero-click queries, SEOs suspect AI fan-out — gaganghotra_ · 2026-09-03
- Two Definitions Of Model Distillation: Raw CoT Jailbreak Vs Task Output Training — JoshPurtell · 2026-09-03
- Gemini 3.8 Flash scores 69.9% on Cursor bench at $2.38 per task — _philschmid · 2026-09-03
- GPT-Astra rumored to be a GPT-4-scale leap as Altman teases 'significant step forwards' — DavidmComfort · 2026-09-03