harnessengineering llm ai agents ai/agenticai

Context

Problem: Long-running agent loops keep accumulating files, commands, and conversation history until the context window becomes overloaded. Performance degrades even before the hard token limit important details and instructions get lost due to **context rot** and eventually the task can fail completely when the limit is reached.

Solution: Treat context as a limited budget, not unlimited storage. Manage it deliberately by compacting/removing less-important context and storing important information externally in memory files, keeping the active context small and useful.

1. Know what you are spending on

Track token usage and understand what is filling the context window. In coding agents, tool outputs—file contents, command logs, and search results—usually consume far more context than prompts or conversation.

The easiest ways to reduce usage are:

  • Truncate large tool outputs and clearly indicate when truncation happened.
  • Read only relevant portions of files using offsets and limits.
  • Let the model fetch information when needed instead of injecting everything upfront, as in retrieval-augmented context selection.

Measure context usage first, then reduce the biggest source of waste—usually tool results.

2. Compaction: forgetting well

When a session gets close to its context limit, compaction replaces older conversation history with a much shorter summary.

flowchart LR
    subgraph before [array at 80% full]
        S1[system] --- O["old rounds<br/>(120k tokens)"] --- R["recent rounds<br/>(20k tokens)"]
    end
    before -->|summarize old rounds| after
    subgraph after [fresh array]
        S2[system] --- SUM["summary<br/>(2k tokens)"] --- R2["recent rounds<br/>(20k, verbatim)"]
    end

Compaction has two important properties:

  • It is lossy.
  • It should happen occasionally, not continuously.
    • One large reset is cheaper and simpler than repeatedly trimming small pieces of context.

Compaction is a Cache Reset, and that is fine, coz its rare.

Compaction deliberately sacrifices old detail to recover working space while preserving the information most important for continuing the task.

3. Memory files: remembering outside the array

Some information should survive beyond one session. Instead of keeping it permanently in the context window, store it in memory files such as CLAUDE.md.

These files can contain durable project knowledge such as:

  • build commands,
  • project structure,
  • coding conventions,
  • warnings,
  • important decisions.

The useful analogy is:

Context array = RAM
Files = disk

Memory files can also be maintained by the model itself. For example, if the user says a project uses pnpm, the model can add that fact to the project’s memory file so future sessions inherit it.

Two important design rules:

  • Inject memory as user/project context, rather than changing the system prompt.
  • Keep memory files small. Use them as an index pointing to larger documents that the model can fetch when needed.

Persistent knowledge belongs outside the active context window, but even persistent memory must be kept compact.

4. The budget mindset

Treat the model’s context window as a limited resource, not storage you can fill indefinitely.

A nominal ~200k-token window may have significantly less space where attention remains consistently reliable. That budget is divided among:

AreaHow to control it
System prompt and toolsKeep them lean; avoid loading rarely needed tools
Memory filesKeep them short and point to deeper documents
Conversation and tool resultsTruncate, fetch on demand, and compact
Free headroomPreserve space for upcoming reasoning and tool use
Eventually, some tasks genuinely require more information than one context window can handle. At that point, trimming alone is not enough—you need to distribute the work elsewhere.

Effective agents actively manage context as a budget, preserving enough free capacity for the model to keep reasoning well.


Subagents

Tldr

Subagents let an agent spend lots of temporary context elsewhere and bring back only the small piece of information worth keeping.

1. The failure and the patch

Some tasks require a huge amount of exploration but produce only a tiny useful answer. If the main agent performs all that exploration itself, thousands of tokens of searches, file reads, and dead ends remain in its context and contribute to context rot.

The solution is a subagent: a separate agent loop with its own fresh context. It completes one task, returns only its final answer to the parent, and then its context is discarded.

Move context-heavy work into a disposable context and keep only the useful conclusion.

2. Subagents are a context tool, not an org chart

Subagents are often described as teams of specialist agents, but their main engineering purpose is simpler:

**They protect the main agent's working memory.**

The parent keeps its important context—task, plan, and decisions—while the child spends its own context budget investigating something. The child might use 50,000 tokens, but the parent may receive only a 200-token summary.

This does not necessarily save API cost. The child still consumes tokens. What it saves is the parent’s limited attention and context capacity.

flowchart TD
    subgraph parent [main agent's array — stays clean]
        P1[task, plan, decisions] --> P2["tool_use: agent('find the retry logic')"]
        P2 --> P3["tool_result: 'Retries live in src/net/backoff.ts,<br/>exponential, used by fetchWithRetry...'"]
        P3 --> P4[work continues, 200 tokens heavier]
    end
    P2 -.spawns.-> C
    subgraph C [subagent's array — disposable]
        C1[fresh array: instructions + the one task]
        C1 --> C2[30 rounds of grep/read/grep...]
        C2 --> C3[final text answer]
    end
    C3 -.only this returns.-> P3

Subagents trade extra computation for cleaner main-agent context.

3. Three structural facts about subagents

The child knows nothing

A fresh subagent does not automatically know:

  • the conversation,
  • the user’s goals,
  • the parent’s plan,
  • previous decisions.

Therefore, the parent must give it a self-contained task description explaining what to do, where to look, and what to return.

Delegation quality depends heavily on the quality of the task brief.

The parent sees only the report

The parent does not normally receive all the searches and reasoning performed by the child. It only receives the final result.

This keeps context clean, but also makes errors harder to inspect. Good subagent reports should therefore include evidence such as file paths, line numbers, or test results.

Hide the exploration, but preserve enough evidence to verify the conclusion.

Results return like any other tool result

From the parent’s perspective, a subagent is simply another tool call:

task sent → subagent works → report returned

This means normal tool practices still apply, including clear failures and error handling.

Architecturally, a subagent can be treated as a tool backed by another agent loop.

4. What to delegate

The best rule is:

Delegate work with large intermediate context but a small final result.

Good examples include:

  • Search and exploration: locating where something is implemented.
  • Verification: running large test suites and summarizing failures.
  • Research: reading several documents and extracting conclusions.
  • Parallel independent tasks: reviewing several unrelated files simultaneously.

These tasks generate lots of temporary information that the parent does not need to retain.

Delegate work where most of the information produced is disposable.

5. What not to delegate

Some work benefits from staying in the main loop.

Poor candidates include:

  • Tasks requiring accumulated judgment. A fresh child does not share the parent’s full understanding of the session.
  • Tiny lookups. Starting another agent loop has overhead; a simple search may be cheaper directly.
  • Long chains of dependent work. Repeated handoffs cause information loss.

The chapter argues for a relatively shallow structure:

one coordinator + disposable workers

rather than deeply nested agent hierarchies.

Keep context-sensitive decision-making centralized and delegate bounded investigations.

6. Production notes

Real systems add features around the same basic idea.

Named agent types

Different subagents can receive different:

  • system prompts,
  • roles,
  • tool permissions.

For example, an exploration agent might be allowed to read files but not modify them.

Different models for different jobs

Simple investigation can be assigned to cheaper models, while harder reasoning stays with a stronger model.

Background execution

Slow subagents can run independently instead of blocking the parent.

Messaging running children

Some systems allow follow-up instructions to an already-running child.

Forked subagents

Not every child must start empty. A forked child can receive a copy of the parent’s current context.

The trade-off is:

Fresh child: cheaper context, less inherited knowledge
Forked child: more inherited knowledge, more context cost

Production systems mainly add routing, permissions, concurrency, and context-sharing options around the same core mechanism.

7. Writing good subagent tasks

A useful mental model is to write the task as if you were assigning work to:

a contractor who has no access to your previous conversations.

A good task should explain:

  • what needs to be done,
  • where to investigate,
  • relevant background,
  • what counts as completion,
  • what evidence to provide,
  • what the final response should contain.

Many apparent “multi-agent coordination” failures are really failures to provide enough context in the delegation.

Clear task specifications are more important than elaborate agent coordination.

This separation between working context and durable memory also appears in agent memory architecture.