harnessengineering llm ai agents ai/agenticai
Core Idea
RAG is deciding what text enters the model’s context. Every RAG variant is ultimately a different answer to: Which text should be selected, and how should it be selected?
The Core Problem
The model does not automatically know:
- your company wiki
- your private codebase
- yesterday’s tickets
- internal documents
- newly created information
And these sources are often too large to paste into the context wholesale.
So the harness must select relevant information before or during reasoning.
large knowledge source -> selection -> small relevant subset -> model contextThat selection process is RAG.
Classic RAG
Classic RAG follows this pattern:
question
↓
harness searches
↓
top relevant chunks
↓
chunks + question enter context
↓
model answersThe important detail:
The harness retrieves before the model reasons.
Example:
def rag_answer(question):
chunks = search(question, top_k=5)
context = "\n\n".join(
chunk.text for chunk in chunks
)
messages = [{
"role": "user",
"content":
f"Answer using this context:\n"
f"{context}\n\n"
f"Question: {question}"
}]
return call_llm(messages)Conceptually:
search + context injection + one model callThere does not need to be an agent loop at all.
The Search Machinery
Most of the intimidating RAG vocabulary lives inside:
search(question)Embeddings
Embeddings convert text into vectors so semantically similar text lands near similar text.
Example:
query:
"How do I get my money back?"
can retrieve:
"Refunds may be requested within 30 days."even though the wording differs.
Think of embeddings as grep for meaning.
Chunking
Documents are divided into smaller retrievable units.
document
↓
chunk 1
chunk 2
chunk 3
chunk 4Retrieval then selects chunks rather than entire documents.
The chunking strategy affects retrieval quality because a chunk should ideally contain enough context to be useful without becoming unnecessarily large.
Vector Database
A vector database stores embeddings and efficiently finds nearby vectors.
query embedding -> nearest-neighbour search -> similar document chunksThis is fundamentally a search infrastructure component, not part of the model itself.
Hybrid Search
Vector search is strong at semantic similarity.
Keyword search is often stronger at:
- exact identifiers
- filenames
- product IDs
- unusual terminology
- rare tokens
So production retrieval often combines both:
keyword search + vector search -> merge / rank -> final candidatesThe RAG Zoo, Decoded
| Name | What it really means |
|---|---|
| Naive RAG | Vector search → paste top-k results |
| Hybrid RAG | Combine keyword + semantic search |
| Re-ranking | Retrieve broadly, then use a better ranker |
| Graph RAG | Search entities and relationships |
| Corrective RAG | Detect weak retrieval and try again |
| Self-RAG | Model evaluates whether retrieval is sufficient |
| Agentic RAG | Model controls retrieval through tools |
The names sound like separate architectures, but most differences are simply:
index choice + ranking choice + retry logicLearn the retrieval knobs, not the taxonomy.
Push vs Pull
This is the chapter’s main distinction.
Push — Classic RAG / Harness Retrieval
The harness guesses what the model will need **before reasoning starts**.
flowchart LR Q[Question] --> S[Harness searches] S --> C[Top chunks] C --> A[Context] A --> M[Model answers]
The body chooses the evidence.
Advantages
- cheap, fast, predictable
- often one search + one LLM call
- easy to operate at high volume
Weakness
The retrieval query is based only on the original question.
If the harness retrieves the wrong material, the model may be confidently grounded in the wrong documents.
Pull — Agentic Retrieval
Instead of injecting everything up front, expose retrieval as a tool. (search_docs(query))
Now the model decides when and what to search.
flowchart TD Q[Question] --> M1[Model reasons] M1 -->|search_docs| S1[Search results] S1 --> M2[Model reads] M2 -->|refined search| S2[Better results] S2 --> M3[Answer]
The LLM (brain) controls the reading.
Example:
Question:
Compare our current refund policy
with the policy before the rebrand.
Model:
"I need the current policy."
→ search_docs("current refund policy")
Model:
"Now I need the previous version."
→ search_docs("refund policy before rebrand")
Model:
compares bothThis is difficult for one-shot retrieval and natural for a loop.
Push vs Pull Trade-off
Pull wins on quality
The model can:
- formulate its own searches
- inspect results
- notice missing evidence
- reformulate queries
- perform multi-hop retrieval
The retrieval process gains a feedback loop.
search -> inspect -> reason -> search againPush wins on cost and latency
Classic RAG can be:
1 search + 1 model callAgentic retrieval might require:
search → model → search → model → search → modelThat costs more tokens, time, and compute.
Design Rule
Pull when reasoning should steer the reading. Push when the reading is predictable.
Push fits
- FAQ bots
- clean support knowledge bases
- predictable single-hop questions
- batch pipelines
- strict latency requirements
Pull fits
- investigation
- research
- ambiguous questions
- multi-hop reasoning
- coding agents
- large heterogeneous knowledge stores
Push + Pull Together
These patterns are not mutually exclusive. A strong architecture can combine both:
obvious context --> push into initial request
additional uncertain context --> pull via tools as neededFor a coding agent:
memory file = push
grep / search / read_file = pullThis hybrid architecture often provides a good cost-quality balance.
Deep Retrieval as a Subagent
Some retrieval tasks consume many rounds and produce a relatively small final result.
Example:
Parent agent
↓
spawn research agent
↓
search
read
search
compare
verify
↓
cited summary
↓
Parent keeps only summaryThis prevents the parent context from filling with retrieval debris.
Fork high-volume work that leaves behind a low-volume result.
Deep-research systems often approximate this architecture.
The Bigger Lesson
RAG appears externally as a specialized AI discipline. From the harness perspective, it collapses into familiar primitives:
stateless context
+
limited context budget
+
search
+
context injection
+
optional tool loop
+
optional subagentThe only substantially new component is the search index. This makes RAG a memory-and-retrieval layer for the agent architecture, not a separate kind of model. Everything else is context engineering.