ai ai/agenticai llm ai/books/designingllmapplication
Tldr
RAG turns an LLM from a system that relies primarily on what it memorized into a system that can dynamically retrieve, select, refine, and use external information when solving a task.
Core idea: RAG augments an LLM with external data by retrieving relevant information at query time and giving it to the model as context. It is especially useful when the model needs private, recent, long-tail, or otherwise unavailable knowledge.
- RAG gives LLMs access to external knowledge without requiring that knowledge to be stored in model parameters.
- Long-tail, private, and recent information are major reasons to use RAG.
- Retrieval quality is the central bottleneck of a RAG system.
- A production RAG pipeline may include:
Rewrite β Retrieve β Rerank β Refine β Insert β Generate β Verify- BM25 and hybrid search remain important even in modern embedding-based systems.
- Reranking improves precision after an initial high-recall retrieval stage.
- Refinement reduces noise and context length through summarization, filtering, or Chain-of-Note.
- Active retrieval can interleave retrieval with generation when additional information is needed.
- RAG can also provide long-term memory, dynamic few-shot examples, tool selection, and training-time retrieval.
- RAG has important limitations: retrieval failures, surface-level reasoning, contradictory evidence, and latency.
- RAG, long context, and fine-tuning are not mutually exclusive; the appropriate choice depends on the task, data, cost, latency, and desired behavior.
- In many knowledge-intensive applications, RAG should be considered before fine-tuning.
1. Why RAG?
LLMs struggle to reliably memorize facts that appear rarely in their training data. Increasing model size helps, but long-tail knowledge remains difficult to memorize. RAG avoids requiring the model to memorize everything by retrieving the needed information from an external data store.
Main reasons to use RAG
-
Private/proprietary data β access information that was not part of pre-training.
-
Reduced hallucination risk β ground answers in retrieved source material and enable citations.
-
Recent information β retrieve information that appeared after the modelβs knowledge cutoff.
-
Long-tail knowledge β answer questions about entities and concepts that occur rarely in training data.
2. Common RAG Scenarios
-
External knowledge retrieval β fill knowledge gaps and reduce hallucinations.
-
Context/history retrieval β retrieve relevant parts of long conversations or sessions.
-
In-context example retrieval β dynamically select useful few-shot examples for the current query.
-
Tool-related retrieval β retrieve tool descriptions, API documentation, and information needed for tool selection.
3. When Should the Model Retrieve?
RAG should not necessarily be used for every query.
Retrieval decisions can depend on:
-
How frequently the entity appears in the modelβs training data.
-
Whether the query requires up-to-date information.
-
Whether the model has been continually trained or memory-tuned on the relevant knowledge.
-
The latency/cost introduced by retrieval.
For very large LLMs, dynamically choosing between parametric memory and retrieval can improve responsiveness. For smaller models, particularly models around 7B parameters or below, RAG is generally more beneficial than relying solely on internal memory.
4. The RAG Pipeline
A production RAG system can be represented as:
Rewrite β Retrieve β Rerank β Refine β Insert β Generate
A verify step can also be applied to outputs from different stages. The retrieve and generate stages are fundamental; the other stages are optional depending on performance and latency requirements.
Rewrite
The user query may use vocabulary different from that used in the documents. Query rewriting attempts to bridge this gap and improve retrieval.
Techniques include:
-
Query expansion with synonyms and related terms.
-
Pseudo-Relevance Feedback (PRF) β retrieve initial documents, extract salient terms, and add them to the query.
-
LLM-based approaches such as Query2doc and HyDE.
-
Query decomposition for complex tasks.
-
Rewriting queries into SQL or equivalent queries for structured databases.
A major risk is topic drift, where rewriting causes the query to move away from the original intent.
Retrieve
Retrieval is the most critical bottleneck in RAG. If the correct information is not retrieved, even an excellent LLM cannot produce the correct answer from that information. Retrieval should therefore emphasize recall.
Important approaches:
-
Embedding-based retrieval β useful for semantic similarity.
-
BM25/keyword retrieval β remains a strong baseline.
-
Hybrid search β combines keyword and embedding-based retrieval.
-
Metadata filtering β restricts retrieval using attributes such as topic or other stored metadata.
Advanced retrieval
-
Generative retrieval: the LLM generates document IDs directly. It is most appropriate for relatively small, low-redundancy, or well-categorized collections.
-
Tightly-coupled retrievers: retrieval is optimized using feedback from the LLM so that retrieved text is useful for generating the correct answer.
-
GraphRAG: builds a knowledge graph from entities and relationships, then uses hierarchical clustering and summaries to answer questions involving relationships and high-level themes. Its major drawback is the substantial computation required to construct the graph.
Rerank
Initial retrieval normally prioritizes recall, so the returned documents may contain irrelevant results.
A reranker sorts the retrieved candidates by relevance and improves precision.
Common approaches include:
-
Cross-encoder relevance models.
-
ColBERT and other late-interaction models.
-
Query likelihood models.
-
LLM-based pointwise, pairwise, and listwise ranking.
Pairwise ranking can be particularly effective because documents are directly compared, although it requires many comparisons.
Refine
Retrieved documents may be too long, noisy, or poorly phrased for the LLM.
The refine stage can:
-
Shorten retrieved content.
-
Rephrase it.
-
Filter irrelevant material.
-
Produce summaries.
-
Generate Chain-of-Note (CoN) representations.
Summarization
-
Extractive: selects important sentences from the original text; generally more faithful.
-
Abstractive: generates a new summary; potentially more useful but carries a greater hallucination risk.
The goal of RAG summarization is not necessarily human readability. The summary should instead help the downstream LLM produce the correct answer.
Chain-of-Note
CoN creates notes explaining:
-
What each retrieved passage says.
-
Whether it directly answers the query.
-
Whether it only provides useful context.
-
Whether it is irrelevant.
-
Whether the retrieved evidence is sufficient to answer the question.
This is especially useful when retrieval contains noise or when the correct answer may not exist in the retrieved documents. It can help the LLM recognize when the correct response is βI donβt know.β
Insert
The refined context must be arranged inside the LLM prompt.
Possible strategies include:
-
Concatenating retrieved documents.
-
Feeding documents separately and combining results.
-
Ordering documents according to relevance.
-
Placing less relevant documents toward the middle of a very long context, since models can recall information at the beginning and end more effectively than information in the middle.
Generate
The LLM uses the query plus retrieved context to generate the final answer.
Generation can be:
-
One-shot: retrieve context first, then generate.
-
Interleaved/active retrieval: generate part of an answer, retrieve additional information when needed, and continue generating.
FLARE is an example of active retrieval. It can trigger retrieval when the model predicts that additional information is needed.
Ground-truth citations are also an important part of generation because they connect generated claims to their supporting sources.
5. RAG Beyond Knowledge Retrieval
Memory Management
RAG can act as an external memory system for an LLM.
Instead of placing an entire conversation into the context window, the system can retrieve only relevant historical information. This resembles virtual memory in an operating system: frequently needed information is brought into the directly accessible context while other information remains in external storage.
This can support:
-
Long-running conversations.
-
Personalization.
-
Persistent preferences.
-
Access to historical interactions.
Recursive summarization is another option, but summarization is lossy and can remove important nuances such as tone.
Selecting Few-Shot Examples
RAG can dynamically retrieve the best few-shot examples for a query instead of using the same examples for every input.
LLM-R improves this idea by using LLM feedback to identify examples that increase the probability of producing the correct answer.
RAG for Model Training
RAG is not limited to inference. It can also participate in pre-training and fine-tuning.
REALM (Retrieval-Augmented Language Model) combines a knowledge retriever with a knowledge-augmented encoder. The retriever learns to find external documents that help the model predict masked tokens during training.
6. Limitations of RAG
RAG is powerful but does not eliminate fundamental LLM limitations.
Key problems
-
Surface-level dependence: the LLM may rely on retrieved snippets without developing deeper understanding.
-
Retrieval bottleneck: bad retrieval leads to bad answers regardless of generator quality.
-
Contradictory information: retrieved content can conflict with the modelβs internal knowledge.
-
Latency: multiple retrieval, reranking, and refinement stages add overhead, so inference optimization can matter.
Contradictory information is particularly difficult because the model does not have direct access to ground truth. The chapter discusses the RECALL benchmark for evaluating robustness to counterfactual information.
7. RAG vs. Long Context
Long-context LLMs can process much larger prompts, reducing the need to retrieve only small portions of a dataset.
However:
-
Large contexts can contain distracting information.
-
Real-world documents often contain related rather than completely unrelated text.
-
Retrieval can be cheaper and faster than placing everything into the context window.
-
Long-context models can be especially useful for analyzing genuinely long documents.
The chapter recommends empirical evaluation rather than assuming that either approach is universally better.
Mental model:
Long context = keep more information directly accessible.
RAG = selectively bring relevant information into the accessible context.
8. RAG vs. Fine-Tuning
Both approaches can incorporate external knowledge, but they do so differently.
| RAG | Fine-tuning |
|---|---|
| Adds knowledge at inference time | Changes model parameters |
| Knowledge can be updated easily | Updating requires additional training |
| Usually less expensive | Training can be expensive |
| Good for knowledge-intensive tasks | Good for domain/task specialization |
| Provides direct access to source material | Can learn patterns and relationships from training data |
The chapter cites research showing that RAG consistently outperformed fine-tuning on knowledge-intensive tasks. However, the approaches can be combined: a model can be fine-tuned for a specialized domain or desired style and then augmented with RAG for external knowledge.
RAG Crash Course for Beginners
How to Build a Scalable RAG System for AI Apps (Full Architecture)
Retrieval Augmented Generation (RAG) Explained: Embedding, Sentence BERT, Vector Database (HNSW)
RAG for Knowledge Intensive NLP Tasks Paper