Unverified: third-party Scry claims 71.8% vs Gemini's 66.1% on DeepSearchQA

AaronBergman18 · x · 2026-09-06

An X account (XyraSinclair, retweeted by AaronBergman18) claims its deep-search product Scry beat Google's Gemini Deep Search Agent on DeepSearchQA, scoring 71.8% against 66.1%.

DeepSearchQA is a 900-prompt arXiv benchmark from a Google team for evaluating deep research agents on multi-step information-seeking across 17 fields. Tasks are causal chains testing fragmented-information collation, deduplication/entity resolution, and stopping-criteria reasoning in open search spaces. The paper finds even state-of-the-art models struggle to balance recall with precision, with failure modes like premature stopping and hedging.

Note: the comparison is a third-party claim, not independently verified.

Original post →

More from Models

Models channel →