ScholarCatalyst: 184 authors label a benchmark for retrieving research-inspiring papers

Sohyeon Kim · hf · 2026-10-03

ScholarCatalyst tests whether AI can find the prior papers a new research question needs. 184 lead authors of 207 recent CS papers labeled which candidates did or could have advanced their projects, with rationales. Results: agentic search (0.42 R@20) underperforms embedding retrieval (0.48), and even a Claude Fable 5.1-based agent reaches only 0.51 — highlighting the gap between AI and expert scientific intuition.

Related event: ScholarCatalyst Benchmark Tests AI Research Taste(4 posts)→

Original post →

More from Research

Research channel →