Grok 4.6 ties Claude Opus 5 on finance diligence bench at ~$0.84/task
karinanguyen · x · 2026-08-19
Benchmark results for Grok 4.6 on deep financial research:
- #2 on DiligenceBench finance harness at 52–53%, effectively tied with Claude Opus 5; Sonnet 5 trails at 46.2%
- Notably different behavior: Grok averaged 41 tool calls per task vs 22 for Opus, and 3,293 SEC filing searches vs 486 — broader search helped it find more evidence and follow instructions closely, while Opus was more efficient and sometimes more nuanced
- Cheaper too: $0.84/task vs $1.02 for Opus 5 despite far more searching
- Grok 4.6 also leads the General Qualitative category on Vals' Finance Agent Benchmark v2 (FAB v2), consistent with the broad evidence-gathering style DiligenceBench rewards
Related event: Grok 4.6 Ties for Second on DiligenceBench at Low Cost(2 posts)→
More from Models
- Qwen 3.8 Comparison: Smaller Q4 Model Outreasons Larger Q5 — k-r-a-u-s-f-a-d-r · 2026-08-19
- GLM-5.3 ties Kimi K3 with score of 60, weights to be released — ArtificialAnlys · 2026-08-19
- Gemini Image Generation Silently Fails From Hetzner IPs — Network Origin Was the Culprit — dota2dinall · 2026-08-19
- GLM-5.3 Scores 60 on AI Index, Touted as Strongest Chinese Model — teortaxesTex · 2026-08-19
- DFlash 2 available for Qwen 3.8 27B and Muse Glimmer — rerri · 2026-08-19
- Recent Codex update broke subagents; rolling back to 0.142.0 works — chibop1 · 2026-08-19