ai llm ai/agenticai ai/books/anillustratedguidetoaiagents

1. Introduction to Harness Engineering

Core idea

An AI agent is an LLM-based system that perceives an environment, reasons about a goal, and acts through tools. It extends a stateless LLM with memory, tools, planning, reflection, and an appropriate level of autonomy.

What is an AI agent?

Russell and Norvig define an agent as:

Quote

Anything that can be viewed as perceiving its environment through sensors and acting upon that environment through actuators.
— Russell & Norvig, Artificial Intelligence: A Modern Approach

For LLM-backed agents, this maps to:

Agent conceptLLM-agent equivalent
EnvironmentThe digital or physical world, including the user
SensorsText, images, audio, video, and environment state
ActuatorsTools, APIs, browsers, shells, and code interpreters
Agent programA reasoning LLM plus memory, tools, planning, and reflection

From LLM to agent

A basic LLM is a stateless text-in/text-out function. It tokenizes input and autoregressively predicts one token at a time. By itself, it cannot directly act on the world or retain information between calls.

A reasoning LLM improves on this by spending additional inference-time compute on intermediate reasoning before producing an answer. This helps with:

  • Breaking complex tasks into subproblems with chain-of-thought reasoning
  • Making multi-step decisions
  • Planning and selecting tools
  • Reflecting on errors and revising actions

The progression is:

  1. LLM: Generates text.
  2. Reasoning LLM: Deliberates over difficult problems.
  3. Augmented LLM: Adds memory and tools.
  4. AI agent: Adds planning, reflection, and goal-directed action.
  5. Agentic system: Places the agent in a larger system with autonomy controls, evaluation, and safety mechanisms.

Core components

Memory

Multi-turn conversations expose a vital flaw of LLMs, namely that they’re forgetful entities and do not remember past conversations. They are stateless, which means that information is not persisted across calls.

LLMs are stateless unless previous information is supplied again. Memory can include conversation history, short-term working context, and long-term knowledge. More context is not always better. Excess information can create information overload and reduce decision quality. Context engineering selects and structures the information most useful for the current task.

Tools

Fundamentally, LLMs can be seen as software or functions that, upon receiving input text, process it and then output some text. As text-in/text-out functions, LLMs can only describe or show the intent of taking the action when outputting text.

An LLM cannot execute an action merely by writing something like multiply(5.1, 7.3).

External software must parse the model’s structured output, select the tool, validate its parameters, execute it, and return the result. Tools connect the agent to the environment through search, APIs, browsers, calculators, shells, and coding environments.

Planning and reflection

Planning decomposes an open-ended goal into smaller, actionable steps. The agent executes the plan incrementally and updates it as new information appears.

Reflection evaluates previous actions, identifies mistakes or missing steps, and revises the plan. This creates an iterative loop:

Autonomy and responsible use

Agency exists on a spectrum. An agent may be allowed to choose one action within a fixed workflow, or it may have broad freedom to select tools, create steps, revise plans, and stop when it believes the task is complete.

Useful safeguards include:

  • Human in the loop: A person authorizes, checks, or audits important decisions.
  • Guardrails: Policies and permissions prevent destructive or unexpected actions.
  • Misinformation controls: Verification is needed because LLMs can confidently produce incorrect information.
  • Evaluation: Assess both the final outcome and the trajectory of actions.

Agent evaluation should measure:

  • Outcome: Was the task completed correctly?
  • Trajectory: Were the steps and tool calls efficient and sound?
  • Reliability: Does the agent succeed consistently across repeated runs?
  • Safety: Does it avoid harm, including harm caused by malicious inputs, manipulated data, or its own mistakes?

Common applications

  • Coding: Read codebases, write or modify code, run tests, and fix errors.
  • Deep research: Search and synthesize information from many sources with limited user intervention.
  • Automation: Navigate heterogeneous data and processes, especially where the goal is clear but the exact procedure is not.

Agent specializations

Part II of the book extends the single-agent foundation to:

  • Multi-agent collaboration: Specialized agents coordinate, often under a supervisor agent.
  • Multimodal agents: Agents understand and/or generate text, images, audio, and video.
  • Coding agents: Agents operate in execution environments and can implement, test, and debug software.

The TinyAgent

The book builds a minimal agent harness incrementally. The initial skeleton is intentionally non-functional:

class TinyAgent:
    """A minimal, modular, and educational agent framework."""
 
    def __init__(self):
        self.llm = None       # Chapters 2–3
        self.memory = None    # Chapter 4
        self.tools = None     # Chapter 5
        self.planner = None   # Chapter 6
 
    def run(self, task: str) -> str:
        return self._step(task)
 
    def _step(self, task: str) -> str:
        return f"Received: {task}"
 
    def _execute_action(self, action: str) -> str:
        return f"Executed action: {action}"

The code-based, terminal-oriented harness demonstrates how an agent can be built from first principles without hiding the important mechanisms behind a framework. Agent harnesses may also be terminal-based, personal assistants, hosted products, or UI-based tools.

Key takeaway

The defining feature of an agent is not simply that it uses an LLM. It is that the system can pursue a goal by deciding what to do, interacting with an environment, using feedback, and adapting its approach. Reasoning provides deliberation, memory supplies context, tools enable action, and planning and reflection turn these capabilities into agentic behavior.