RAG (Retrieval-Augmented Generation) 检索增强生成
What is RAG?
RAG (Retrieval-Augmented Generation) is an important solution in the large language model field for addressing factual accuracy problems.
By dynamically retrieving from an external knowledge base, the model can access up-to-date information at inference time, forming a hybrid architecture of "pretrained model + dynamic knowledge base".
This fundamentally addresses the "knowledge cutoff" and "factual hallucination" problems of traditional language models.
Core Concepts of RAG
1. Retrieval
- Retrieve relevant document fragments from an external knowledge base
- Use semantic similarity matching
- Combine the retrieved information with the original question
- Build a rich contextual prompt
- Generate accurate answers based on the augmented context
- Answers can be traced to their sources
- ✅ Improved factual accuracy — reduces model "hallucinations" by retrieving real data
- ✅ Dynamic knowledge updates — the knowledge base can be updated without retraining
- ✅ Strong domain adaptability — quickly adapt to different professional domains by swapping the knowledge base
- ✅ Enhanced explainability — answer references are traceable
- ⚠️ Retrieval quality dependency — retrieval quality directly affects final generation quality
- ⚠️ Increased latency — the retrieval step adds extra computation and I/O overhead
- ⚠️ Knowledge update costs — requires maintaining a high-quality, timely-updated knowledge base
- ⚠️ Context length limits — retrieved content may exceed the model's context window
- Intelligent retrieval — precise document retrieval based on semantic similarity
- Dynamic augmentation — combines retrieved information with user queries in real time
- Precise generation — generates accurate, relevant, and traceable answers grounded in the augmented context
- Enterprise knowledge base Q&A systems
- Intelligent customer service assistants
- Academic research support
- Medical diagnosis support
- Legal consultation services
2. Augmentation
3. Generation
Advantages of RAG
Challenges
RAG System Architecture
Core Modules
1. User interface — receives questions and displays results 2. Orchestrator — coordinates modules and manages the overall workflow 3. Retrieval module — retrieves relevant document fragments for the user query (semantic retrieval, BM25 algorithm, vector similarity) 4. Knowledge base — stores and manages external knowledge sources (vector databases, Elasticsearch, FAISS) 5. Context builder — combines retrieval results with the user question into complete context 6. Large language model — generates the final answer based on the augmented context
RAG Workflow
1. User inputs a question 2. Retrieve relevant documents — fetch relevant document fragments from the knowledge base 3. Augment the context — combine retrieved documents with the original question 4. Generate the final answer — produce an accurate answer based on the augmented context