ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
To study if AI can match human scientific intuition in identifying prior research that advances a project, the authors build ScholarCatalyst using annotations and detailed rationales from 184 lead authors of 207 recent computer science papers. They introduce a retrieval task based on author-provided judgments using literature available when projects began. Experimental results show that agentic search achieves 0.42 Recall@20 versus 0.48 for embedding retrieval, while an agent built on Claude Fable 5.1 reaches only 0.51 R@20. These findings highlight the need for new training recipes that provide models with expert search intuition.
Agentic search does no better than embedding retrieval with 0.42 vs. 0.48 Recall@20.
An agent built on Claude Fable 5.1 reaches only 0.51 R@20 on the ScholarCatalyst retrieval task.
ScholarCatalyst is built from judgments by 184 lead authors of 207 recent computer science papers.