Query Expansion with LLMs: Searching Better by Saying More (Jina AI, Feb 2025)
Source: Jina AI blog
This post is a structured digest of Jina AI's February 2025 article on using large language models for query expansion in information retrieval.
Key points
- Core idea: User queries are typically short and ambiguous, while relevant documents use richer vocabulary. An LLM can rewrite, paraphrase, and enrich a query — adding synonyms, domain terms, and decomposing multi-intent queries into sub-queries — to improve recall for both sparse (BM25) and dense retrievers.
- Typical pipeline: input representation → expansion/decomposition → retrieval → reranking → generation, with early stopping and compute budgets controlling latency and cost.
- Positioning in the IR landscape: The work sits between classic cascaded retrieval (recall → rerank → generate) and newer agentic paradigms, where the number of retrieval steps and search strategy themselves become learnable decisions.
- Evaluation considerations: Standard metrics (nDCG@10, MRR, Recall@k) plus task success rates; LLM-as-judge evaluation should be cross-validated with human assessment since offline metrics increasingly diverge from online satisfaction.
- Latency & cost: Expansion adds an LLM call per query; mitigate with caching of popular queries, cascaded designs, small/distilled models for expansion, and p99 budget controls.
- Quality: Verify that offline gains translate to online CTR/satisfaction via interleaving experiments and manual audits.
- Safety: Open-ended query rewriting can introduce poisoning, bias, or hallucinated terms; consider source allowlists and output filtering.
- Data: Synthetic training data must guard against knowledge leakage and distribution shift.
- Reliable evaluation of expansion quality
- Latency/cost trade-offs in production
- Cross-lingual and multimodal generalization
- Deeper integration with knowledge graphs and structured databases
Engineering takeaways
Open problems
Context and related reading
The post cross-references related items: ChatGPT Search clickstream analysis, Phi-3 as a search relevance judge, Netflix's foundation model for personalized recommendation, and dynamic filtering for web search accuracy.
> Note: This is a metadata-based digest of the Jina AI article. Specific quantitative results should be verified against the original source.
Glossary
| Term | Meaning | |------|---------| | IR | Information Retrieval | | RAG | Retrieval-Augmented Generation | | nDCG | Normalized Discounted Cumulative Gain, a ranking quality metric | | Agentic Search | Modeling search as sequential decision-making and tool use | | Gen-IR | Generative Information Retrieval |