A new benchmark, ScholarCatalyst, has been developed to evaluate artificial intelligence systems' ability to identify research papers that inspire new scientific work. The benchmark uses annotations from 184 lead authors of computer science papers who identified which prior works influenced their projects.

The ScholarCatalyst benchmark aims to quantify a critical skill in scientific progress: the ability to connect existing ideas with new research questions. Developed by researchers including Sohyeon Kim and Graham Neubig, the benchmark comprises 207 computer science papers and approximately 191,000 candidate documents. It includes 894 research questions, each with author-provided judgments on which papers could have advanced their work, along with detailed rationales. This dataset is intended to facilitate the development of AI systems that can assist scientists by pointing to relevant prior research, a task that current AI struggles with compared to human researchers.

The creation of ScholarCatalyst involved an automated pipeline to scale author annotations. Researchers were asked to verify research questions as they stood before their projects' key findings and then label candidate papers that did or could have influenced their work, even if they had not encountered them at the time. This approach grounds the benchmark in the firsthand knowledge of researchers about their own projects, offering a unique perspective on how scientific inspiration is formed and disseminated.

The benchmark is designed to support a retrieval task where AI systems must identify inspiring papers given an initial research question. This differs from existing benchmarks that may focus on general reasoning or novelty judgment of ideas. ScholarCatalyst specifically targets the nuanced process of scientific hypothesis formulation by focusing on the "inspiration retrieval" component. The researchers envision ScholarCatalyst as a step towards creating "scientific agents" capable of taking a half-formed idea and directing researchers to the necessary prior literature.

The development of ScholarCatalyst addresses a gap in current AI evaluation, where systems excel at solving defined problems but lag behind human scientists in sensing the relevance of buried ideas within a vast research archive. By using author-provided rationales, the benchmark provides a rich source of information for training and evaluating AI models on this complex cognitive task. The dataset and associated code are available for use by the research community.