A Survey of Model Architectures in Information Retrieval
Overview
- Paper: A Survey of Model Architectures in Information Retrieval
- Authors: Zhichao Xu, Fengran Mo, Zhiqi Huang, Crystina Zhang, Puxuan Yu, Bei Wang, et al. (8 authors total)
- Published: 2025-02-20
- Type: Survey
- A unified perspective that places scattered related work into a comparable framework.
- Clear decomposition of method components: representation learning, retrievers, re-rankers, planners, generators, and feedback mechanisms.
- Reproducible benchmarks, datasets, and taxonomy tables that lower the entry cost for follow-up researchers.
- Discussion of interfaces with LLM tool use, reinforcement learning, and multi-agent collaboration, including paths from research prototypes to industrial systems.
- Explicit open problems: evaluation trustworthiness, latency and cost, hallucination and safety, and cross-lingual/multimodal extension.
- Dense retrieval: high recall, low latency; suits first-stage retrieval.
- Late interaction (e.g., ColBERT): higher accuracy but larger indexes.
- Generative IR: directly "generates" documents via tokens or docids, simplifying cascades.
- Agentic search: models search as sequential decision-making, supporting multi-hop reasoning and self-reflection.
- 2019–2021: BERT re-ranking and DPR establish neural retrieval foundations.
- 2022–2023: RAG and FreshLLMs drive retrieval-generation fusion.
- 2024 onward: conversational/agentic search and Gen-RecSys surge.
- 2025–2026: RL-trained search agents, Deep Research, and GraphRAG become new growth areas.
- Datasets: MS MARCO, BEIR, Natural Questions, domain corpora, recommendation sets.
- Metrics: nDCG@10, MRR, Recall@k, Hit@k, human preference, task success rate, latency and token cost.
- Baselines: BM25, dense retrieval, cross-encoder re-ranking, retrieval-free LLMs, commercial search APIs.
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search
- A Survey of Conversational Search (arXiv 2410.15576, Oct 2024)
- A Survey of LLM-Empowered Agents for Recommendation (arXiv 2503.05659)
- A Survey on Knowledge-Oriented Retrieval-Augmented Generation (arXiv 2503.10677)
- Original paper: A Survey of Model Architectures in Information Retrieval. arXiv:2502.14822
Original Abstract
> The period from 2019 to the present marks one of the most significant paradigm shifts in information retrieval (IR) and natural language processing (NLP), culminating in the emergence of powerful large language models (LLMs) from 2022 onward. Methods based on pretrained encoder-only architectures (e.g., BERT) as well as decoder-only generative LLMs have outperformed many earlier approaches, demonstrating particularly strong performance in zero-shot scenarios and complex reasoning tasks. This survey examines the evolution of model architectures in IR, with a focus on two key aspects: backbone models for feature extraction and end-to-end system architectures for relevance estimation. To maintain analytical clarity, we deliberately separate architectural design from training methodologies, enabling a focused examination of structural innovations in IR systems. We trace the progression from traditional term-based retrieval models to modern neural approaches, highlighting the transformative impact of transformer-based architectures and subsequent LLM developments. The survey concludes with a forward-looking discussion of open challenges and emerging research directions, including architectural optimization for efficiency and scalability, robust handling of multimodal and multilingual data, and adaptation to novel application domains such as autonomous search agents, which may represent the next paradigm in IR.
Key Points
Research Motivation
Large-scale search, recommendation, and personalization systems face long-standing challenges in efficiency, scalability, and user-intent understanding. Traditional pipelines separate retrieval, ranking, and generation, which struggles to meet LLM-era demands for natural-language interaction, multi-hop reasoning, and real-time knowledge. This survey systematically organizes theory and practice at this intersection.
Contributions
Taxonomy and Technical Routes
| Dimension | Subcategory | Representative Ideas | Strengths | Limitations | |---|---|---|---|---| | Modeling paradigm | Discriminative / generative retrieval | Bi-encoders, cross-encoders, DSI, GPT-based indexing | Mature, scalable | Semantic drift, update cost | | LLM integration | RAG / Agent / Tool-use | Retrieval augmentation, search agents, API calls | Flexible, interpretable | Latency, error propagation | | Optimization objectives | Relevance / diversity / freshness | Multi-objective LTR, RLHF, online learning | Business-aligned | Scarce annotations | | Evaluation | Offline / Online / Human | nDCG, MRR, LLM-as-judge, A/B testing | Comparable | Gap vs. real satisfaction |
Four main lines are compared:
Timeline of Evolution
Evaluation Paradigm
Typical benchmarks and trends covered:
For exact quantitative results, consult the original PDF; the survey recommends cross-checking numbers before citing.
Main Takeaways for Search / Rec / Personalization
1. Architecture: cascaded retrieval + re-ranking + generation remains mainstream, but the agentic paradigm is making retrieval count and strategy itself learnable. 2. Data: high-quality instruction data and click/session logs are both critical; synthetic data needs protection against leakage and distribution shift. 3. Evaluation: the gap between offline metrics and online satisfaction is widening; LLM-as-judge needs cross-validation with human assessment. 4. Product: latency, cost, interpretability, and safety are hard constraints for industrial deployment—academic benchmarks alone are insufficient.
Open Problems and Future Directions
Noted gaps include lack of unified benchmarks, non-reproducible private data, LLM evaluation bias, and safety/cost constraints for agentic systems. Future work includes finer-grained process supervision, joint retrieval-reasoning training, enterprise metadata governance, and multimodal/cross-lingual consistency.
Related Entries
Glossary
| Term | Meaning | |---|---| | IR | Information Retrieval | | RAG | Retrieval-Augmented Generation | | LTR | Learning to Rank | | nDCG | Normalized Discounted Cumulative Gain | | Agentic Search | Modeling search as sequential decisions and tool calls | | Gen-IR | Generative Information Retrieval |