Deep research and web retrieval benchmarks are stale — Reddit calls for real-environment evals
Lucky_Creme_5208 · reddit · 2026-09-26
- A Reddit user argues most benchmarks are either saturated or tested under artificially restricted conditions — e.g., AA Omniscience restricts tool access.
- Deep research, wide web search and fact retrieval are among the most common chatbot use cases, yet dedicated benchmarks are scarce, and the few that exist have no official leaderboard and haven't been updated in years.
- The post calls for a benchmark run in real environments — with tool access, internet search, and real harnesses like Claude or ChatGPT Work — and asks the community whether one exists.
Related event: Community Calls for Real-World Search Benchmarks as Existing Ones Saturate(2 posts)→
More from Research
- Contrastive World Models Learns World Models in Latent Space Without Pixel Prediction — burny_tech · 2026-09-27
- Rethinking on-policy distillation: researchers propose OLIVE, letting students learn from teacher continuations — May_F1_ · 2026-09-27
- CoRL 2026 workshop on continually self-improving robots opens call for papers, due Sep 28 — PeterStone_TX · 2026-09-27
- Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs — AccBalanced · 2026-09-27
- Functional Gradient Descent with Adaptive Representations accepted at NeurIPS — CatAstro_Piyush · 2026-09-27
- Tailored ASR for Japanese speaking assessment cuts mora error rate from 12.3% to 7.1% — tkasasagi · 2026-09-27