Artificial Analysis shows Gemini 3.6 Flash and 3.5 Flash-Lite improve on agentic work
ArtificialAnlys · x · 2026-07-21
Artificial Analysis published the full benchmark set for Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.
- Both models improved on GDPval-AA v2 and AA-Briefcase, two agentic knowledge-work benchmarks.
- Gemini 3.6 Flash reached 1421 Elo on GDPval-AA v2 and 961 Elo on AA-Briefcase.
- Gemini 3.5 Flash-Lite jumped to 1140 Elo on GDPval-AA v2 and 634 Elo on AA-Briefcase, showing a much larger gain than 3.6 Flash relative to its predecessor.
More from Models
- Google launches three new Gemini models, including a cybersecurity system — Polymarket · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22