Unverified: third-party Scry claims 71.8% vs Gemini's 66.1% on DeepSearchQA
AaronBergman18 · x · 2026-09-06
An X account (XyraSinclair, retweeted by AaronBergman18) claims its deep-search product Scry beat Google's Gemini Deep Search Agent on DeepSearchQA, scoring 71.8% against 66.1%.
DeepSearchQA is a 900-prompt arXiv benchmark from a Google team for evaluating deep research agents on multi-step information-seeking across 17 fields. Tasks are causal chains testing fragmented-information collation, deduplication/entity resolution, and stopping-criteria reasoning in open search spaces. The paper finds even state-of-the-art models struggle to balance recall with precision, with failure modes like premature stopping and hedging.
Note: the comparison is a third-party claim, not independently verified.
More from Models
- Rumor: Astra Low to X-High Shows Only 1% Reasoning Difference — BLUECOW009 · 2026-09-06
- Unverified "GPT-6 Astra" rumor goes viral: from next-word predictor to frontier model — DeryaTR_ · 2026-09-06
- "Best model ever, yet I have to babysit it": dev bemoans agentic reliability gap — yacineMTB · 2026-09-06
- Cybersecurity pro says Claude guardrails block legit work, asks about Astra — GVTZ59 · 2026-09-06
- Fable 5.1 drops medical refusals, but its 85% life-sciences claim draws scrutiny — JeremyNguyenPhD · 2026-09-06
- ChatGPT Pro's free usage reset also pushes back your weekly limit reset date — dmpizzle · 2026-09-06