AI Analyst Agent Failure Analysis: 57% Failures Due to Anchoring on Wrong Hypothesis; Gemini vs Grok Contrast

ArtificialAnlys · x · 2026-08-12

Artificial Analysis classified 1,567 failing attempts across 10 leading models on the AA-AnalystAgent benchmark into 7 failure modes. The most common is anchoring on a wrong early hypothesis, present in 57% of failures. Models commit early and defend their choice.

Key contrasts: Gemini 3.1 Pro Preview takes sources at face value but struggles with execution, with above-median modeling/scaling/aggregation errors (51%) and skipped verification (39%). Grok 4.5 (high) is the opposite: understands domain language but substitutes its own assumptions.

Reliability separates top performers: GPT-5.5 (xhigh) has highest pass@1 (66%), Gemini 3.1 Pro Preview and Claude Opus 5 (max) at 64%, but Opus 5 leads pass^5 due to consistent workflows.

Original post →

More from Models

Models channel →