English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DeepWeb-Bench: A Harder Deep Research Benchmark Requiring Massive Cross-Source Evidence and Long-Horizon Derivation

Forum topic · 小凯 · 2026-05-22

Summary

DeepWeb-Bench (arXiv 2505.15982) is a deep research benchmark designed to be substantially harder than existing evaluations for frontier language models. Its difficulty stems from three data properties: massive evidence collection, cross-source reconciliation, and long-horizon multi-step derivation. These map to four capability families—Retrieval, Derivation, Reasoning, and Calibration—with results reported per family. Each reference answer includes source provenance with four disclosure levels and auditable cross-source checks. Evaluations across nine frontier models yield three findings: retrieval is not the bottleneck (retrieval failures account for only 12–14% of errors, while derivation and calibration failures exceed 70%); strong and weak models fail in qualitatively different ways (strong models show incomplete derivations, weak models show hallucinated precision); and models exhibit genuine domain specialization, with cross-model agreement of only ρ=0.61 and per-case divergence up to 18.8 percentage points. The public release includes data, scoring rubrics, and evaluation code.

Overview

  • Field: Machine Learning
  • Authors: Sixiong Xie, Zhuofan Shi, Haiyang Shen
  • Published: 2025-05-20
  • arXiv: 2505.15982

Abstract

Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. However, frontier deep research products score high on existing benchmarks, making it difficult to distinguish their capabilities from current evaluation data alone.

The authors introduce DeepWeb-Bench, a deep research benchmark that is substantially harder than existing benchmarks for the current frontier. Difficulty comes from three properties of the data itself:

1. Massive evidence collection — each task requires gathering large amounts of evidence. 2. Cross-source reconciliation — evidence must be coordinated across sources. 3. Long-horizon multi-step derivation — answers require extended multi-step reasoning.

These three sources of difficulty are represented as four capability families (Retrieval, Derivation, Reasoning, and Calibration), with results reported sliced by family. Every reference answer comes with source provenance at four disclosure levels and usable cross-source checks, making scores easier to audit against the underlying evidence.

Key Findings

Evaluating DeepWeb-Bench on nine frontier models, the paper reports three main findings:

1. Retrieval is not the bottleneck. Retrieval failures account for only 12–14% of errors, while derivation and calibration failures account for over 70%. 2. Strong and weak models fail in qualitatively different ways. Errors from strong models are dominated by incomplete derivations, whereas weak models exhibit hallucinated precision. 3. Models show genuine domain specialization. Cross-model agreement is only ρ=0.61, with per-case divergence reaching 18.8 percentage points.

Resources

The public benchmark release includes data, scoring rubrics, and evaluation code.

--- *Paper: arXiv:2505.15982*

Tags

#deep-research#benchmark#llm-evaluation#retrieval-augmented-generation#web-agents#machine-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620575