English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LLM-Driven Usefulness Judgment for Web Search Evaluation

Forum topic · 小凯 · 2026-07-05

Summary

This April 2025 arXiv paper (arXiv:2504.14401) by Mouly Dewan, Jiqun Liu, Aditya Gautam, and Chirag Shah explores using large language models (LLMs) to produce usefulness judgments for web search evaluation. Traditional search evaluation relies heavily on relevance judgments (e.g., graded relevance for metrics like nDCG), but usefulness—how well a document actually helps a user complete their search task—is harder and more costly to assess with human assessors. The authors investigate whether LLMs can serve as reliable judges of usefulness, examining how well LLM-generated usefulness judgments align with human judgments and how they compare with conventional relevance-based evaluation. The work situates itself in the broader shift toward LLM-as-a-judge evaluation pipelines in information retrieval, touching on issues of evaluation cost, scalability, and the gap between topical relevance and actual task utility. It is relevant to researchers working on search evaluation methodology, LLM-based judging, and the design of user-centric retrieval metrics. Readers should consult the full PDF for detailed experimental setups and quantitative results.

LLM-Driven Usefulness Judgment for Web Search Evaluation

Paper: https://arxiv.org/abs/2504.14401 (arXiv, April 2025)

Authors: Mouly Dewan, Jiqun Liu, Aditya Gautam, Chirag Shah

Overview

This paper examines whether large language models (LLMs) can be used to generate usefulness judgments for web search evaluation. Classical evaluation of search systems depends on human relevance judgments, which are expensive to collect and measure *relevance* rather than *usefulness*—that is, whether a document genuinely helps a user accomplish their underlying task.

Key Points

  • Motivation: Relevance judgments (e.g., graded relevance feeding nDCG-style metrics) do not fully capture whether a result actually serves a user's search task. Usefulness judgments are more user-centric but costly to gather from human assessors.
  • Approach: The authors leverage LLMs as automated judges to assess the usefulness of web documents in the context of a user's information need, following the broader "LLM-as-a-judge" paradigm.
  • Evaluation concern: The work studies the alignment between LLM-generated usefulness judgments and human judgments, and how LLM-based usefulness evaluation compares with traditional relevance-based evaluation.
  • Context: The paper contributes to ongoing efforts to reduce the cost and increase the scalability of search evaluation while shifting metrics closer to actual user benefit.
  • Why It Matters

  • Search evaluation has long been criticized for the gap between offline relevance metrics and real user satisfaction. Automating usefulness judgments with LLMs could make user-centric evaluation practical at scale.
  • It connects to larger trends in IR: generative retrieval, RAG, and agentic search all require evaluation methods that go beyond static relevance labels toward task success and utility.
  • The approach raises open questions typical of LLM-judge pipelines: reliability, bias, cost/latency, and the need for cross-validation against human assessment.
  • Limitations

    As with most LLM-as-judge work, potential concerns include consistency across runs, sensitivity to prompt design, domain and language coverage, and disagreement between LLM and human notions of usefulness. Quantitative details, datasets, and agreement statistics are available in the full paper PDF; readers should consult the original before citing specific numbers.

    References

  • Original paper: LLM-Driven Usefulness Judgment for Web Search Evaluation

Tags

#information-retrieval#llm-as-a-judge#search-evaluation#usefulness-judgment#web-search#relevance-judgment#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208702