← Back to Paper List

ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry

Tianze Xu, Pengrui Lu, Lyumanshan Ye, Xiangkun Hu, Pengfei Liu
Shanghai Jiao Tong University, SIIGAIR
arXiv.org (2025)
Agent Benchmark Factuality RAG

📝 Paper Summary

Benchmark datasets Agentic RAG pipeline AI Research Evaluation
ResearcherBench is the first benchmark designed to evaluate Deep AI Research Systems on their ability to act as genuine research partners for open-ended, frontier scientific questions.
Core Problem
Existing benchmarks evaluate deep research systems on retrieving and synthesizing established knowledge, failing to assess their capacity to generate novel insights for genuinely unsolved, frontier scientific problems.
Why it matters:
  • True research assistance requires deep conceptual understanding and the generation of novel insights, not just information gathering
  • Current automated LLM-as-a-judge criteria fail on cutting-edge research tasks because they lack the extensive domain expertise needed to identify high-value insights
  • Deploying AI on frontier research creates a feedback loop for recursive self-improvement toward Artificial Superintelligence (ASI)
Concrete Example: When asked a cutting-edge question from a laboratory discussion where no definitive answer exists, current evaluation frameworks focus on the breadth of retrieved information rather than assessing if the AI synthesized disparate ideas into a conceptually deep, valuable research insight.
Key Novelty
ResearcherBench: Dual Evaluation for Frontier AI Research
  • Curates 65 authentic, high-value research questions from real-world scenarios like lab discussions and expert interviews across 35 Artificial Intelligence (AI) subjects
  • Employs a dual evaluation framework combining expert-crafted rubrics for insight quality with automated factual verification for citation accuracy
Evaluation Highlights
  • OpenAI Deep Research and Gemini Deep Research significantly outperform other baseline systems on frontier research tasks
  • Leading commercial systems show particular strength in exploratory reasoning for open-ended consulting questions rather than precise technical synthesis
  • Evaluation uncovers a 'high faithfulness, low groundedness' pattern, indicating that explicit citations are accurate but many generated claims lack citation support
Breakthrough Assessment
8/10
Pioneers the evaluation of AI systems as genuine research partners on unsolved scientific frontiers, introducing a rigorous dual-assessment methodology that combines human expertise with automated verification.
×