ScholarCatalyst Benchmark Tests AI Research Taste
CMU, UW and collaborators released ScholarCatalyst, a benchmark where 184 first authors annotated the 'catalyst papers' behind 207 recent projects. Embedding retrieval currently outperforms agentic search on the task.
2026-10-02 ~ 2026-10-03 · 4 related posts
- ScholarCatalyst: a new benchmark testing whether AI can pick research problems like humans — PangWeiKoh · 2026-10-02
- ScholarCatalyst: New Benchmark Shows Agentic Search Loses to Plain Embedding Retrieval — lateinteraction · 2026-10-02
- ScholarCatalyst: a benchmark for 'research taste', labeled by 184 authors — lateinteraction · 2026-10-02
- ScholarCatalyst: 184 authors label a benchmark for retrieving research-inspiring papers — Sohyeon Kim · 2026-10-03