harnessengineering llm ai agents ai/agenticai

Core Idea

The prompt is advice. The body is enforcement.

An agent is probabilistic, and its context can contain untrusted instructions. Therefore, the model cannot also be the safety mechanism. Safety must live in deterministic layers between what the model requests and what the machine actually executes.

Core Problem

There are two major sources of unsafe behaviour:

  1. Ordinary model mistakes
    The model may misunderstand a task and request a destructive action.
  2. Prompt injection
    Tool results, files, web pages, and MCP descriptions can contain malicious instructions that influence the model.

The key architectural rule is:

The model requests actions; the harness decides whether those actions happen.

The Guardrail Stack

Model requests action
        ↓
Layer 0 β€” Is capability available?
        ↓
Layer 1 β€” Deterministic rules
        ↓
Layer 2 β€” Human approval
        ↓
Layer 3 β€” Model-based risk classifier
        ↓
Layer 4 β€” Sandbox / limited blast radius
        ↓
Execute

The layers move from cheap and certain toward more flexible but less reliable mechanisms.

Layer 0 β€” Don’t Give It the Hands

The safest dangerous tool is one that does not exist.

If an agent only needs to read files, do not expose write tools.
If a subagent only summarizes websites, give it fetching capability and nothing else.

No capability = No possible misuse

This is capability minimization.

Installing an MCP server is effectively granting the agent a new set of hands.

Layer 1 β€” Deterministic Reflexes

Before executing a tool request, ordinary code can inspect it.

Useful rules include:

  • protect .env, ~/.ssh, system directories, etc.
  • block --force operations
  • enforce allowlists / denylists
  • limit output sizes
  • prevent edits to files the agent has not read
  • reject edits when the file changed after the agent last saw it

The file-read tracker is especially powerful:

read file
   ↓
remember version
   ↓
model requests edit
   ↓
has file changed?
   ↓
yes β†’ refuse + explain
no  β†’ continue

These rules are cheap and deterministic.

Prompt:
"Don't overwrite stale files"
β‰ˆ probabilistic
 
Harness invariant:
"Reject edits when version changed"
= guaranteed

When rejecting an action, return the reason to the model so it can recover.

Guardrails can simultaneously enforce and steer.

Layer 2 β€” Ask the Human

Some actions cannot be classified safely with simple rules.

For example:

rm -rf build/

could be completely normal or disastrous depending on context.

The harness therefore creates an approval gate.

Unlike an ask_user tool, the model does not choose whether to ask. The harness itself requires approval.

Main danger: approval fatigue

If users must approve everything, they eventually stop evaluating requests carefully.

Ways to reduce this:

  • Permission modes
    • ask for everything
    • auto-approve reads
    • ask for writes
    • ask only for dangerous actions
  • Remembered grants
    • e.g. npm test β†’ always allow
  • Scoped autonomy
    • approve a plan once
    • let the agent execute its normal steps

Every approval can become configuration and reduce future friction.

Layer 3 β€” Cheap Brain Judges Big Brain

Rules cannot classify every possible shell command.

For ambiguous actions, the harness can ask a smaller model:

User request + Proposed action -> Small classifier model -> allow / ask / block

Example:

User:
"Fix the README"
 
Agent:
curl remote-script | bash
 
Classifier:
This action does not match the task.
β†’ escalate

But the classifier is also probabilistic and vulnerable to injection.

Therefore:

Treat classifier output as evidence, not authority.

Good design:

  • verify the classifier’s reasoning
  • bias uncertain cases toward human approval
  • make false-allow errors much harder than false-ask errors
  • never use the classifier as the only safety layer

Layer 4 β€” Limit the Blast Radius

Eventually, prevention may fail.

The final layer assumes something unsafe gets through and limits the damage.

Recoverability

Git makes many coding actions reversible.

bad edit β†’ git restore / checkout β†’ recover

An environment with easy undo needs fewer gates than one where changes are irreversible.

Sandboxing

Run tools inside:

  • containers
  • VMs
  • restricted OS profiles

Restrict:

  • filesystem access
  • network access
  • process capabilities

The sandbox does not care whether the model was mistaken, injected, or confidently wrong.

It enforces boundaries at the operating-system level.

Minimal credentials

Do not give the agent secrets it does not need.

Credential unavailable --> credential cannot be leaked

System Prompts Are Not Access Control

A rule such as:

NEVER delete files outside the project.

is usefulβ€”but it is still only text presented to a probabilistic model.

It can fail because of:

  • prompt injection
  • context dilution
  • sampling variance
  • model mistakes

The proper hierarchy is:

Prompts      β†’ advise
Reflexes     β†’ enforce
Humans       β†’ decide
Classifiers  β†’ screen
Sandboxes    β†’ contain

Anything that must always be true should be enforced outside the model.