DiligenceBench benchmarks equity-research agents, with Muse Spark 1.1 reaching 57.4%
karinanguyen · x · 2026-07-22
DiligenceBench benchmarks equity-research agents
A new benchmark, DiligenceBench, evaluates whether AI agents can do real public-equity research.
- The charted results show Meta Muse Spark 1.1 leading the finance harness at 57.4%, followed by GLM 5.2, Sonnet 4.6, and GPT-5.6 Sol.
- The authors found that stronger models benefit most from generic tools that unlock execution, while weaker models need more opinionated, domain-specific scaffolding.
- For Inkling, a generic sandbox barely helped (20.9% → 22.5%), but the finance harness raised it to 32.8%, with the biggest gain coming from factual accuracy.
- The benchmark also changes the price–performance frontier: most models get both better and cheaper under the finance harness, with GLM 5.2 leading on absolute performance and MiniMax M3 appearing among the efficiency winners.
More from Venture
- Discovery Summit panel says AI is outrunning science funding and lab workflows — allisondman · 2026-07-22
- A billboard claims one company can deliver an AI data center in nine months — BenBajarin · 2026-07-22
- Meitu launches a RMB 100 million challenge for live AI imaging apps — 量子位 · 2026-07-22
- MSA Pairformer beats ESMC 6B while academia still struggles to raise $100K — KevinKaichuang · 2026-07-22
- AI one-person companies are surging as Abacus AI bundles backend, hosting, and payments — bindureddy · 2026-07-22
- Anthropic-Physical Intelligence acquisition rumor spreads as 2026 AI M&A heats up — TechCrunch AI · 2026-07-22