English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Survey: LLM-Empowered Agents for Recommendation and Search Toward Next-Gen Information Retrieval (arXiv 2503.05659)

Forum topic · 小凯 · 2026-07-05

Summary

This post reviews the March 2025 arXiv survey 'A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval' (arXiv:2503.05659) by Yu Zhang, Shutong Qiao, Jiaqi Zhang, Tzu-Heng Lin, Chen Gao, and Yong Li. The survey systematically organizes work at the intersection of large language models, agentic search, and recommendation systems. It provides a unified taxonomy covering modeling paradigms (dense, late-interaction, generative retrieval), LLM integration patterns (RAG, tool use, search agents), optimization objectives, and evaluation protocols. The post outlines the field's evolution from BERT-based reranking and DPR (2019-2021), through RAG-driven retrieval-generation fusion (2022-2023), to the current wave of conversational/agentic search, generative recommender systems, RL-trained search agents, and GraphRAG. It also summarizes benchmark datasets (MS MARCO, BEIR, Natural Questions), common metrics (nDCG, MRR, Recall@k, LLM-as-judge), open problems such as evaluation reliability, latency, hallucination, and safety, plus a practical engineering checklist for deploying LLM-based retrieval and recommendation systems.

Survey: LLM-Empowered Agents for Recommendation and Search (arXiv 2503.05659)

This forum post introduces and analyzes the survey paper "A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval" (March 2025, arXiv).

Paper metadata

| Field | Content | |-------|---------| | Title | A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval | | Authors | Yu Zhang, Shutong Qiao, Jiaqi Zhang, Tzu-Heng Lin, Chen Gao, Yong Li | | Published | March 2025, arXiv | | Link | https://arxiv.org/abs/2503.05659 | | Type | Survey |

Background and motivation

Agentic search has long faced challenges in efficiency, scalability, and user intent understanding. Traditional pipeline approaches split retrieval, ranking, and generation into disjoint stages, which struggles to meet LLM-era demands for natural language interaction, multi-hop reasoning, and up-to-date knowledge. The survey aims to systematically organize the theory and practice of this intersection, covering open-domain information access, enterprise knowledge retrieval, conversational search, and semantic understanding in recommendation systems.

Core contributions

  • A unified perspective that brings scattered prior work into a comparable framework.
  • A clear decomposition of method components: representation learning, retrievers, rerankers, planners, generators, and feedback mechanisms.
  • Reproducible benchmarks, datasets, and taxonomy tables to lower entry barriers for follow-up research.
  • Discussion of interfaces with LLM tool calling, reinforcement learning, and multi-agent collaboration, including paths from research prototypes to industrial systems.
  • Explicit open problems: evaluation trustworthiness, latency and cost, hallucination and safety, and cross-lingual/multimodal extension.
  • Taxonomy

    | Dimension | Subcategories | Representative ideas | Strengths | Limitations | |-----------|---------------|----------------------|-----------|-------------| | Modeling paradigm | Discriminative vs. generative retrieval | Bi-encoders, cross-encoders, DSI | Mature, scalable | Semantic drift, update cost | | LLM integration | RAG / Agent / Tool-use | Retrieval augmentation, search agents, API calls | Flexible, interpretable | Latency, error propagation | | Optimization objectives | Relevance / diversity / freshness | Multi-objective LTR, RLHF, online learning | Business-aligned | Scarce annotations | | Evaluation | Offline / Online / Human | nDCG, MRR, LLM-as-judge, A/B tests | Comparable | Diverges from real satisfaction |

    Main research threads

    The survey contrasts four major lines of IR research:

  • Dense retrieval: high recall and low latency; suited to first-stage retrieval.
  • Late interaction (e.g., ColBERT): higher precision but larger index footprint.
  • Generative IR: "generates" documents directly via tokens or docids, simplifying cascades.
  • Agentic search: models search as sequential decision-making, supporting multi-hop reasoning and self-reflection.
  • Timeline of the field

  • 2019-2021: BERT reranking and DPR establish neural retrieval.
  • 2022-2023: RAG and FreshLLM drive retrieval-generation fusion.
  • 2024: Explosion of conversational/agentic search and Gen-RecSys.
  • 2025-2026: RL-trained search agents, Deep Research, and GraphRAG become new growth fronts.
  • Evaluation practices

    Typical benchmarks and trends covered:

  • Datasets: MS MARCO, BEIR, Natural Questions, domain corpora, public recommendation datasets.
  • Metrics: nDCG@10, MRR, Recall@k, Hit@k, human preference, task success rate, latency, token cost.
  • Baselines: BM25, dense retrieval, cross-encoder reranking, retrieval-free LLMs, commercial search APIs.
  • Precise numbers should be checked against the original PDF tables.

    Key insights

    1. Architecture: cascade retrieval + rerank + generation remains mainstream, but agentic paradigms make "how often and where to retrieve" itself learnable. 2. Data: high-quality instruction data and click/session logs are both critical; synthetic data risks leakage and distribution shift. 3. Evaluation: the gap between offline metrics and online satisfaction is widening; LLM-as-judge should be cross-validated with human evaluation. 4. Products: latency, cost, explainability, and safety are hard constraints for industrial deployment—do not optimize academic benchmarks alone.

    Engineering deployment checklist

    | Item | Question | Suggestion | |------|----------|------------| | Data | PII in training/index? Versioning? | Partitioned indexes, anonymization, rollback-capable embedding versions | | Latency | p99 budget? Retrieval steps? | Cascade + early stopping, hot-query caching, async reranking | | Quality | Do offline gains translate to online CTR/satisfaction? | Interleaving experiments, human audits, citation verification | | Safety | Poisoning/bias via open retrieval? | Source whitelists, adversarial detection, output filtering | | Cost | Token and GPU cost per query? | Small-model routing, distillation, hybrid sparse+dense |

    Open problems and future directions

    The field lacks unified benchmarks; private data limits reproducibility; LLM evaluation is biased; and agentic systems face safety and cost constraints. Future directions include finer-grained process supervision, joint retrieval-reasoning training, enterprise metadata governance, and multimodal/cross-lingual consistency, as well as deeper fusion with knowledge graphs and causal/fairness constraints for recommendation.

    Related entries

  • A Comprehensive Survey on Reinforcement Learning-based Agentic Search (arXiv:2510.16724)
  • A Survey of Conversational Search (arXiv:2410.15576)
  • A Survey of Model Architectures in Information Retrieval (arXiv:2502.14822)
  • A Survey on Knowledge-Oriented Retrieval-Augmented Generation (arXiv:2503.10677)
  • Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions

Tags

#llm-agents#information-retrieval#recommender-systems#agentic-search#rag#survey#generative-retrieval#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208971