ScholarCatalyst benchmark shows frontier models can't spot which prior papers inspire new research

CatAstro_Piyush · x · 2026-10-05

Researchers from Stanford, CMU, MIT and others built ScholarCatalyst, a retrieval benchmark testing whether AI can identify which prior papers would advance a new research problem. 184 lead authors of 207 recent CS papers labeled which candidates did or could have advanced their projects, with rationales. Results: agentic search on Claude Fable 5.1 hits only 0.51 Recall@20 vs 0.48 for plain embedding retrieval, even though the agent calls that same retriever as a tool. The gap suggests models lack scientists' intuition for navigating broad literature, motivating new training recipes for scientific agents.

Original post →

More from Research

Research channel →