Gemini 3.6 Flash lands at 1421 on a real-world task leaderboard

teortaxesTex · x · 2026-07-22

A post reacting to Google’s Gemini 3.6 Flash says the model shows decent progress versus 3.5, but also argues that Google will need a larger model to keep up.

The attached image is a GDPval-AA v2 leaderboard, which measures performance on real-world work tasks and is anchored to a human baseline of 1,000. In the chart:

The quoted Google reply says AA mostly covers reasoning benchmarks, which is why that score did not move much, while the company focused on agentic use cases for real-world tasks.

Related event: Gemini 3.6 Flash Benchmarks Lag Behind Predecessor(30 posts)→

Original post →

More from Models

Models channel →