ScholarCatalyst benchmark shows frontier models can't spot which prior papers inspire new research
CatAstro_Piyush · x · 2026-10-05
Researchers from Stanford, CMU, MIT and others built ScholarCatalyst, a retrieval benchmark testing whether AI can identify which prior papers would advance a new research problem. 184 lead authors of 207 recent CS papers labeled which candidates did or could have advanced their projects, with rationales. Results: agentic search on Claude Fable 5.1 hits only 0.51 Recall@20 vs 0.48 for plain embedding retrieval, even though the agent calls that same retriever as a tool. The gap suggests models lack scientists' intuition for navigating broad literature, motivating new training recipes for scientific agents.
More from Research
- Generative end-to-end ad retrieval at Douyin headlines weekly IR papers roundup — _reachsumit · 2026-10-05
- The top 50 most-cited AI researchers, where one paper carries 278k citations — ksprdk · 2026-10-05
- Dev post-trains Yandex's 80B model from scratch: first SFT round underfits at 5M tokens — jjusko20 · 2026-10-05
- PDMD distillation cuts video diffusion to 2-4 NFEs, ComfyUI weights released — Total-Resort-3120 · 2026-10-05
- University of Tokyo grows living human skin on a robotic finger that self-heals in a week — bennash · 2026-10-05
- INRIA Researcher on the Crisis of Math Research in the Age of AI — BachFrancis · 2026-10-05