Gemini 3.8 Flash Hits 59 on AA Intelligence Index but Burns More Output Tokens
haider1 · x · 2026-09-03
Per @haider1's observation of the AA Intelligence Index, Gemini 3.8 Flash scores 59, a big jump from Flash 3.7.
The odd part: it uses more output tokens per task than the GPT-5.6 family, Opus 5, Fable 5, and even Fable 5.1. Intelligence improved a lot, but token efficiency looks rough, meaning real-world costs may run higher.
More from Models
- 3.8 Flash hits the Pareto frontier: 5x cheaper, 2x faster than Opus 5 at similar intelligence — davidtsong · 2026-09-03
- Pedro Domingos blasts OpenAI: looped transformers aren't more opaque, CoT transparency is fiction — pmddomingos · 2026-09-03
- Google Search Console social properties show bizarre zero-click queries, SEOs suspect AI fan-out — gaganghotra_ · 2026-09-03
- Two Definitions Of Model Distillation: Raw CoT Jailbreak Vs Task Output Training — JoshPurtell · 2026-09-03
- Gemini 3.8 Flash scores 69.9% on Cursor bench at $2.38 per task — _philschmid · 2026-09-03
- GPT-Astra rumored to be a GPT-4-scale leap as Altman teases 'significant step forwards' — DavidmComfort · 2026-09-03