ScholarCatalyst: New Benchmark Shows Agentic Search Loses to Plain Embedding Retrieval
lateinteraction · x · 2026-10-02
A team including Chelsea Finn and Yejin Choi released ScholarCatalyst, a benchmark for retrieving papers that inspire new research. Key findings:
- An automated pipeline had 184 lead authors of 207 recent CS papers label which candidates did or could have advanced their projects, each with a rationale
- Task: given an initial research question, retrieve those papers using only literature available when the project began
- Agentic search hit just 0.42 Recall@20 vs. 0.48 for embedding retrieval — despite calling that same retriever as a tool; even an agent on Claude Fable 5.1 reached only 0.51
- The authors argue new training recipes are needed to give models expert intuition for broad-corpus search
- Paper, code, dataset and website are all publicly available
Related event: ScholarCatalyst Benchmark Measures Research Taste(3 posts)→
More from Research
- Marin's 535B MoE hero run hits ~27% MFU on 11 NVL72 racks: expert parallelism deep dive — dlwh · 2026-10-03
- SWE-sweep benchmark tests whether AI agents can find bugs before users hit them — klieret · 2026-10-02
- Harvard-led team unveils brain imaging 60x faster than fMRI, tracking activity in ~100ms — melnykowycz · 2026-10-02
- RExBench: best coding agent implements research extensions only 33% of the time — najoungkim · 2026-10-02
- Retrieval-augmented episodic memory narrows LLMs' lexical frequency gap in syntax tests — najoungkim · 2026-10-02
- EMPIRIC teaches robots missing physics as code, solving all 25 tasks where baselines get 14-16 — tomssilver · 2026-10-02