Gemini 3.6 Flash lands at 1421 on a real-world task leaderboard
teortaxesTex · x · 2026-07-22
A post reacting to Google’s Gemini 3.6 Flash says the model shows decent progress versus 3.5, but also argues that Google will need a larger model to keep up.
The attached image is a GDPval-AA v2 leaderboard, which measures performance on real-world work tasks and is anchored to a human baseline of 1,000. In the chart:
- Claude 4.5 (with fallback): 1748
- GPT-5.6 (max): 1736
- Kimi K3: 1679
- Claude Sonnet 5 (max): 1607
- Claude Opus 4.8 (max): 1594
- GPT-5.6 (max): 1584
- Grok 4.5 (high): 1535
- GLM-5.2 (max): 1514
- GPT-5.5 (high): 1490
- Gemini 3.6 Flash: 1421
The quoted Google reply says AA mostly covers reasoning benchmarks, which is why that score did not move much, while the company focused on agentic use cases for real-world tasks.
Related event: Gemini 3.6 Flash Benchmarks Lag Behind Predecessor(30 posts)→
More from Models
- Kimi K3 tops DesignArena 3D Design with a 1450 Elo score — tokenbender · 2026-07-22
- A January 2025 pretraining cutoff is “mind-boggling” for a small-compute model — _arohan_ · 2026-07-22
- Qwen3.8 preview brings Alibaba’s open-weight push to enterprise distribution — krishnan · 2026-07-22
- Google reportedly dropped a couple of models today, with more launches teased — bindureddy · 2026-07-22
- Local models face a single-shot HTML flight simulator test across six runs — JLeonsarmiento · 2026-07-22
- poolside/Laguna-S-2.1-GGUF starts trending on Hugging Face — poolside · 2026-07-22