DiligenceBench benchmarks equity-research agents, with Muse Spark 1.1 reaching 57.4%
karinanguyen · x · 2026-07-22
DiligenceBench benchmarks equity-research agents
A new benchmark, DiligenceBench, evaluates whether AI agents can do real public-equity research.
- The charted results show Meta Muse Spark 1.1 leading the finance harness at 57.4%, followed by GLM 5.2, Sonnet 4.6, and GPT-5.6 Sol.
- The authors found that stronger models benefit most from generic tools that unlock execution, while weaker models need more opinionated, domain-specific scaffolding.
- For Inkling, a generic sandbox barely helped (20.9% → 22.5%), but the finance harness raised it to 32.8%, with the biggest gain coming from factual accuracy.
- The benchmark also changes the price–performance frontier: most models get both better and cheaper under the finance harness, with GLM 5.2 leading on absolute performance and MiniMax M3 appearing among the efficiency winners.
More from Venture
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11
- Moonshot's annualized revenue jumped from $300M to $1B in two months after Kimi K3 — Hesamation · 2026-09-11
- Mid-market companies' AI SEO bottleneck is ops execution, not strategy, says SEO practitioner — gaganghotra_ · 2026-09-11
- A YouTuber with 1.5M followers paid this indie maker for a consulting call — tibo_maker · 2026-09-11
- 71% of people have never used generative AI — the bubble argument for microsaas — iamaliveix · 2026-09-11
- Glean grew from $100M to $300M ARR in roughly fifteen months — yogthinks · 2026-09-11