harnessengineering llm ai agents ai/agenticai
Core Idea
The prompt is advice. The body is enforcement.
An agent is probabilistic, and its context can contain untrusted instructions. Therefore, the model cannot also be the safety mechanism. Safety must live in deterministic layers between what the model requests and what the machine actually executes.
Core Problem
There are two major sources of unsafe behaviour:
- Ordinary model mistakes
The model may misunderstand a task and request a destructive action. - Prompt injection
Tool results, files, web pages, and MCP descriptions can contain malicious instructions that influence the model.
The key architectural rule is:
The model requests actions; the harness decides whether those actions happen.
The Guardrail Stack
Model requests action
β
Layer 0 β Is capability available?
β
Layer 1 β Deterministic rules
β
Layer 2 β Human approval
β
Layer 3 β Model-based risk classifier
β
Layer 4 β Sandbox / limited blast radius
β
ExecuteThe layers move from cheap and certain toward more flexible but less reliable mechanisms.
Layer 0 β Donβt Give It the Hands
The safest dangerous tool is one that does not exist.
If an agent only needs to read files, do not expose write tools.
If a subagent only summarizes websites, give it fetching capability and nothing else.
No capability = No possible misuseThis is capability minimization.
Installing an MCP server is effectively granting the agent a new set of hands.
Layer 1 β Deterministic Reflexes
Before executing a tool request, ordinary code can inspect it.
Useful rules include:
- protect
.env,~/.ssh, system directories, etc. - block
--forceoperations - enforce allowlists / denylists
- limit output sizes
- prevent edits to files the agent has not read
- reject edits when the file changed after the agent last saw it
The file-read tracker is especially powerful:
read file
β
remember version
β
model requests edit
β
has file changed?
β
yes β refuse + explain
no β continueThese rules are cheap and deterministic.
Prompt:
"Don't overwrite stale files"
β probabilistic
Harness invariant:
"Reject edits when version changed"
= guaranteedWhen rejecting an action, return the reason to the model so it can recover.
Guardrails can simultaneously enforce and steer.
Layer 2 β Ask the Human
Some actions cannot be classified safely with simple rules.
For example:
rm -rf build/could be completely normal or disastrous depending on context.
The harness therefore creates an approval gate.
Unlike an ask_user tool, the model does not choose whether to ask. The harness itself requires approval.
Main danger: approval fatigue
If users must approve everything, they eventually stop evaluating requests carefully.
Ways to reduce this:
- Permission modes
- ask for everything
- auto-approve reads
- ask for writes
- ask only for dangerous actions
- Remembered grants
- e.g.
npm testβ always allow
- e.g.
- Scoped autonomy
- approve a plan once
- let the agent execute its normal steps
Every approval can become configuration and reduce future friction.
Layer 3 β Cheap Brain Judges Big Brain
Rules cannot classify every possible shell command.
For ambiguous actions, the harness can ask a smaller model:
User request + Proposed action -> Small classifier model -> allow / ask / blockExample:
User:
"Fix the README"
Agent:
curl remote-script | bash
Classifier:
This action does not match the task.
β escalateBut the classifier is also probabilistic and vulnerable to injection.
Therefore:
Treat classifier output as evidence, not authority.
Good design:
- verify the classifierβs reasoning
- bias uncertain cases toward human approval
- make false-allow errors much harder than false-ask errors
- never use the classifier as the only safety layer
Layer 4 β Limit the Blast Radius
Eventually, prevention may fail.
The final layer assumes something unsafe gets through and limits the damage.
Recoverability
Git makes many coding actions reversible.
bad edit β git restore / checkout β recoverAn environment with easy undo needs fewer gates than one where changes are irreversible.
Sandboxing
Run tools inside:
- containers
- VMs
- restricted OS profiles
Restrict:
- filesystem access
- network access
- process capabilities
The sandbox does not care whether the model was mistaken, injected, or confidently wrong.
It enforces boundaries at the operating-system level.
Minimal credentials
Do not give the agent secrets it does not need.
Credential unavailable --> credential cannot be leakedSystem Prompts Are Not Access Control
A rule such as:
NEVER delete files outside the project.is usefulβbut it is still only text presented to a probabilistic model.
It can fail because of:
- prompt injection
- context dilution
- sampling variance
- model mistakes
The proper hierarchy is:
Prompts β advise
Reflexes β enforce
Humans β decide
Classifiers β screen
Sandboxes β containAnything that must always be true should be enforced outside the model.