ScholarCatalyst: 184 authors label a benchmark for retrieving research-inspiring papers
Sohyeon Kim · hf · 2026-10-03
ScholarCatalyst tests whether AI can find the prior papers a new research question needs. 184 lead authors of 207 recent CS papers labeled which candidates did or could have advanced their projects, with rationales. Results: agentic search (0.42 R@20) underperforms embedding retrieval (0.48), and even a Claude Fable 5.1-based agent reaches only 0.51 — highlighting the gap between AI and expert scientific intuition.
Related event: ScholarCatalyst Benchmark Tests AI Research Taste(4 posts)→
More from Research
- Nathan Lambert launches Trillium Labs, a nonprofit for open frontier AI science — morgymcg · 2026-10-03
- AC2: actor-critic with action chunking enables partial rollouts, trains faster than GRPO — ZeYanjie · 2026-10-03
- KAIST's World Observer gives world models movable panoramic eyes to track unseen regions — Scobleizer · 2026-10-03
- Ofir Press defends new bug-finding benchmark: training the behavior is fine, test-set contamination is not — OfirPress · 2026-10-03
- CAS Five-Year Plan gives AI for Science its own chapter, setting up a US-vs-China metascience bet — teortaxesTex · 2026-10-03
- Draft standard for autonomous labs emerges as AI-driven science heats up — Afinetheorem · 2026-10-03