Gemini 3.7 Flash tops new AA-AnalystAgent quantitative analysis benchmark
_philschmid · x · 2026-08-19
Gemini 3.7 Flash achieved the #1 spot on the new AA-AnalystAgent benchmark by Artificial Analysis. The benchmark tests AI agents on 80 real-world quantitative analysis tasks across 14 domains like finance and healthcare, simulating the daily work of business and data analysts. Models run inside an isolated Python 3.12 sandbox using spreadsheets and documents. The ranking metric is pass@5, requiring correct answers across all 5 independent runs to emphasize reliability over one-shot success.
Related event: Gemini 3.7 Flash Tops AA-AnalystAgent Benchmark(2 posts)→
More from Models
- Harmony tags confuse ChatGPT, causing incorrect DeepSeek explanation — mitsuhiko · 2026-08-20
- Unsloth Releases Qwen3.8-27B GGUFs with 10% Higher Accuracy — danielhanchen · 2026-08-20
- Hugging Face releases SmolLM3 mid-training checkpoint amid 200x efficiency debate — eliebakouch · 2026-08-20
- Claude 5.6 Sol Ultra with computer use called a qualitative leap like Opus 4.5 with MCP — curious_vii · 2026-08-20
- Gemini 3.7 Flash Test: 350 tok/s Speed, Mixed Coding Results — haider1 · 2026-08-20
- AntLing open-sources Ling-3.0 models using WSM to replace LR decay — AcanthisittaOk1699 · 2026-08-19