English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Forum topic · 小凯 · 2026-07-05

Summary

DeepResearch Bench (arXiv:2506.11763) is a benchmark proposed by researchers including Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao to evaluate deep research agents powered by large language models. The benchmark assesses agents along two dimensions: the quality of generated research reports and the effectiveness of their web search behavior. For report evaluation, it introduces RACE (Reference-based Adaptive Critique-Evaluate), a framework that leverages LLM-as-judge against a curated set of reference materials to score reports across multiple dimensions. For search capability, it provides a collection of real-world research tasks requiring complex web information retrieval. The work addresses the gap of standardized evaluation for agentic search systems that combine iterative retrieval, planning, and long-form generation, and positions itself within a growing line of research on deep research agents and LLM-based scientific agents. Related surveys and benchmarks are cross-referenced for readers building retrieval-generation-evaluation pipelines.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

arXiv: https://arxiv.org/abs/2506.11763 Authors / Affiliations: Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao Type: Academic paper Section: Deep Research, Evaluation of Search Engines

Overview

DeepResearch Bench is a comprehensive benchmark for evaluating deep research agents — LLM-based systems that perform open-ended research by combining multi-step web search, planning, and long-form report generation. The paper was posted to arXiv in June 2025.

The benchmark targets two core capabilities:

1. Report generation quality — evaluating the depth, breadth, accuracy, and presentation of long-form research reports produced by agents. 2. Web search capability — assessing how effectively an agent retrieves and selects information from the live web for complex research tasks.

Evaluation Methodology

For report quality assessment, the paper introduces RACE (Reference-based Adaptive Critique-Evaluate), a framework that uses LLM-as-judge to compare generated reports against a curated set of reference documents, producing task-specific, criterion-based scores. This aims to improve the reliability of automatic evaluation compared with generic LLM scoring.

For search ability, the benchmark provides research tasks grounded in real-world information needs, requiring iterative querying, source selection, and synthesis.

Key points

Tags

#deep-research-agents#benchmark#llm-evaluation#agentic-search#rag#llm-as-judge#information-retrieval#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208560