Gemini 3.7 Flash tops benchmark, beating flagship frontier models
mtizard · x · 2026-08-20
Google's Gemini 3.7 Flash achieved a score of 60.0% on the Artificial Analysis AA-AnalystAgent benchmark, taking the #1 spot.
Rankings:
- Gemini 3.7 Flash: 60.0%
- Claude Opus 5: 53.8%
- GPT-5.5: 50.0%
- Claude Fable 5: 48.8%
- GPT-5.6 Sol: 47.5%
- Grok 4.6: 41.3%
This result shows a Flash-tier model outperforming several flagship frontier models, suggesting Google AI may not be as far behind as perceived.
Related event: Gemini 3.7 Flash Tops AA-AnalystAgent Benchmark(4 posts)→
More from Models
- Qwen3.8-27B-OBLITERATED Released: A Red-Teamed Uncensored Model — OBLITERATUS · 2026-08-20
- GPT-5.6 Ultra mode shows little coding gain over Extra High in testing — techartist_ · 2026-08-20
- Comment: GPT-5.6 Sol notorious for searching online instead of solving — zainhas · 2026-08-20
- DeepSeek V5 Suspected Testing in the Wild; Claude Code Gets Concise Mode — WorldofAI · 2026-08-20
- SpaceXAI Swaps Land for 69 Acres, Pledges $40M for Public Safety Facilities — chrisgrayson · 2026-08-20
- Users miss cold, objective AI: old prompts to cut filler talk no longer work — Wargaming123A · 2026-08-20