Gemini 3.8 Flash Hits 59 on AA Intelligence Index but Burns More Output Tokens

haider1 · x · 2026-09-03

Per @haider1's observation of the AA Intelligence Index, Gemini 3.8 Flash scores 59, a big jump from Flash 3.7.

The odd part: it uses more output tokens per task than the GPT-5.6 family, Opus 5, Fable 5, and even Fable 5.1. Intelligence improved a lot, but token efficiency looks rough, meaning real-world costs may run higher.

Related event: Gemini 3.8 Flash review: smarter but pricier, lands on the price-performance Pareto frontier(9 posts)→

Original post →

More from Models

Models channel →