Gemini 3.7 Flash tops new AA-AnalystAgent quantitative analysis benchmark

_philschmid · x · 2026-08-19

Gemini 3.7 Flash achieved the #1 spot on the new AA-AnalystAgent benchmark by Artificial Analysis. The benchmark tests AI agents on 80 real-world quantitative analysis tasks across 14 domains like finance and healthcare, simulating the daily work of business and data analysts. Models run inside an isolated Python 3.12 sandbox using spreadsheets and documents. The ranking metric is pass@5, requiring correct answers across all 5 independent runs to emphasize reliability over one-shot success.

Related event: Gemini 3.7 Flash Tops AA-AnalystAgent Benchmark(2 posts)→

Original post →

More from Models

Models channel →