English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Don't Be Fooled by AI's 'Omniscience': Why Your AI Assistant Is Still a Terrible Investigator — DeepWeb-Bench Explained

Forum topic · QianXun · 2026-05-22

Summary

A new benchmark called DeepWeb-Bench, published on arXiv in May 2026 by researchers at Peking University (arXiv ID 2605.15830), exposes the core weakness of today's deep research AI agents: they excel at finding information but fail at reasoning over it. The benchmark requires massive cross-source evidence gathering, reconciling conflicting data across sources, and long-horizon multi-step derivation. Evaluating nine top-tier models including GPT-5 series and Claude 4, the study found that retrieval failures account for only 12–14% of total errors, while derivation and calibration failures exceed 70%. Weaker models tend to exhibit 'fake precision' — confidently fabricating exact-looking numbers when data is missing — while top models fail at incomplete derivation, gathering all the right evidence but making errors in the final calculation steps. The paper also highlights a persistent 'black box': it remains unclear how models internally weigh conflicting evidence, whether by source authority, frequency, or opaque heuristics. The findings suggest that AI evaluation should shift from testing search and memory capabilities toward testing rigorous analysis, evidence reconciliation, and calibrated reasoning — a warning that even polished, thousand-word AI research reports may rest on broken logical chains.

Don't Be Fooled by AI's 'Omniscience': Why Your AI Assistant Is Still a Terrible Investigator

| Attribute | Details | | :--- | :--- | | Paper | DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation | | Authors | Sixiong Xie, Zhuofan Shi, Haiyang Shen et al. (Peking University) | | arXiv ID | 2605.15830 (May 2026) | | Field | Deep Research Agents, complex reasoning, evaluation benchmarks | | Keywords | massive evidence, cross-source reconciliation, fake precision, long-horizon derivation, calibration |

Imagine hiring a private investigator to look into a multinational company's financial misconduct, and the investigator returns in half an hour with fifty beautifully formatted spreadsheets. You'd think their efficiency was remarkable.

But then you ask: "Based on these reports, what was the company's per-vehicle profit last year?"

They may stare at the reports for a long time and then hand you a number that looks extremely precise but is logically wrong. Or worse — since the number isn't in the reports, they might simply invent a figure with two decimal places to bluff their way through.

This predicament — excellent at finding material, terrible at drawing conclusions — is exactly the fatal weakness that all top AI models face on the road to "deep research."

In May 2026, researchers at Peking University published a major paper on arXiv: DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation.

Through a highly challenging benchmark, they unmasked AI's facade of being "broadly knowledgeable but shallowly capable."

Searching Isn't Reasoning: The Three Hurdles of Deep Research

Today's AI already navigates web search with ease. Ask it a celebrity's birthday and it answers instantly.

But ask it: "Compare Tesla's and BYD's per-vehicle R&D spending in Q3 2024, and explain why their gross margins differ?" — that's real "deep research." The paper identifies three mountains such tasks must cross:

1. Massive evidence gathering: extracting fragments from dozens or even hundreds of web pages. 2. Cross-source reconciliation: if a financial report is in RMB and a news article is in USD, the model must unify units and filter out noise. 3. Long-horizon multi-step derivation: the answer is buried in that pile of data, but it takes complex arithmetic and nested logic to extract it.

A Striking Finding: Search Is Solved, Logic Is Not

The team evaluated nine top models, including the GPT-5 series and Claude 4, and reached a conclusion that overturns conventional assumptions:

Retrieval failures account for only 12–14% of total errors.

This shows that current AI agents are already excellent "research assistants" — they can locate nearly all the raw evidence you need.

The real disaster zone is "derivation" and "calibration," with failure rates above 70%.

Models can be grouped into two failure patterns:

  • Weaker models: addicted to "fake precision." When they can't find exact data, they project extreme confidence and fabricate a number specific to two decimal places. This kind of "confident nonsense" is the most typical failure mode of weak models.
  • Top models: defeated by "incomplete derivation." They collect all the right ingredients but fall apart while cooking. For example, when computing "BYD's net profit per vehicle," they correctly identify total revenue, total profit, and sales volume — but in the final step, they use "total gross profit" where "automotive-segment gross profit" was required.

The Derivation Black Box No One Has Opened

Although the paper precisely diagnoses the disease, a deep "black box" remains when examining the underlying reasoning:

How do models implicitly "balance" conflicting evidence?

When one page says A and another says B, does the model's internal attention mechanism judge based on source authority, information frequency, or some ineffable "feel" for language? Although the paper introduces the concept of "calibration," there is still no transparent mathematical explanation of how models weigh evidence during derivation.

Summary

The mark of intelligence is weaving truth's fabric from tangled strands of information.

The paper's message is that the focus of AI progress has shifted.

The arrival of DeepWeb-Bench signals that the evaluation community is no longer satisfied with testing AI's "memory" or "search ability." True deep research demands that AI transform from a "diligent porter" into a "rigorous analyst."

The next time an AI hands you a ten-thousand-word research report, don't just marvel at the speed.

Interrogate it like a skeptical advisor — check that final step where it reaches its conclusion. Because behind that seemingly certain answer may lie a "hollow chain" of broken logic.

Truth often lives not at the top of a search box, but is born from the deep reorganization of evidence. That is the ultimate warning about evidence and logic that deep-research evaluation in 2026 delivers. This article is an English adaptation of a Chinese forum post on zhichai.net; the benchmark's own contribution claims, including its name, authorship, and arXiv ID, are reproduced here as stated in the source post.

Tags

#deep-research#ai-benchmarks#llm-reasoning#deepweb-bench#calibration#hallucination#evidence-reconciliation#evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620580