harnessengineering llm ai agents ai/agenticai
Core Idea
A reliable production harness retries only transient failures, streams and accumulates complete responses, never trusts partial output, manages token-based concurrency, and keeps retries beneath the agent loop so side effects cannot accidentally run twice.
1. Introduction: Why this matters in production
The earlier harness design assumed that model API calls always succeed and treated streaming as mostly a presentation concern. In real production systems, neither assumption holds.
Requests fail regularly, long calls can exceed timeouts, and users notice poor responsiveness when they wait on a spinner. This appendix focuses on the operational details that turn a demo into a reliable service: handling failures, retrying correctly, streaming responses, and accounting for LLM-specific behavior.
2. The kinds of failure
The first important distinction is between temporary failures and permanent failures.
| Failure | What to do |
|---|---|
| Server overload / 500-series errors | Retry with backoff |
| Rate limit / 429 | Respect retry-after, then retry |
| Timeout or dropped connection | Retry safely |
| Context too long | Do not retry; compact context or fail upward |
| Invalid request / 400 | Do not retry; fix the request |
| Authentication / billing / 401–403 | Stop and inform the user/operator |
The key rule is:
Only retry errors that have a reasonable chance of succeeding on the next attempt.
Retrying a temporary server failure makes sense. Retrying a malformed request just wastes money and hides the underlying bug.
For retryable failures, the standard strategy is:
- exponential backoff,
- add jitter so parallel agents do not retry simultaneously,
- honor
retry-afterheaders, - limit the number of attempts,
- report the actual final error instead of a vague “retries exhausted” message.
3. How streaming actually works
With streaming enabled, the model provider keeps the HTTP connection open and sends many small events rather than waiting to return one complete JSON response.
These usually arrive through Server-Sent Events (SSE).
A response might arrive as a sequence such as:
message_start → block starts → text fragments → block stops → usage information → message finishes.
The harness must accumulate all the fragments until it reconstructs the same complete assistant message it would have received from a normal non-streaming request.
This is why streaming is described as a presentation-layer detail:
It changes how the response travels, not the final message stored in conversation history.
Once accumulation finishes, the rest of the harness can operate exactly as if the request had never been streamed.
4. Tool calls also stream
Streaming is more complicated when the model is producing tool calls.
Tool arguments may arrive as incomplete JSON fragments, for example:
{"path": "main.py"}Those fragments cannot safely be interpreted until the entire tool call has arrived.
Therefore:
Never execute a partially streamed tool call.
The harness should wait until the tool-call block is complete before parsing or executing it.
Streaming tool calls mainly improves the user interface—for example, the UI can show that the model is preparing a tool call—but it does not allow tool execution to safely begin earlier.
5. Streaming becomes operationally necessary
Streaming is not merely a UI optimization.
Large requests with:
- long contexts,
- long generated outputs,
- extensive reasoning,
can exceed provider limits for ordinary non-streaming requests.
As production agents grow more complex, systems often end up streaming almost every request and internally accumulating the result, even when no user interface needs to display individual tokens.
Streaming also gives an important operational signal: time to first token.
If a connection remains open but produces nothing for a long time, the system can distinguish that from a model that is actively reasoning and continuously emitting events.
So streaming becomes both:
- a transport mechanism,
- and a health/liveness signal.
6. What is specific to LLM APIs?
The appendix identifies four important differences between ordinary API retry logic and LLM systems.
6.1 Retries are safe because the model API is stateless
A model API request does not permanently change server-side conversational state.
Therefore, retrying the same request does not create the normal API problem of wondering whether the previous operation “half succeeded.”
However, this safety disappears once your own harness starts executing tools.
For example, if a model response triggers a command and you retry the entire agent-loop round, the command might execute twice.
So retries should occur inside call_llm, below the agent loop.
Do not rerun the entire loop round.
6.2 Rate limits are about tokens, not just request count
LLM providers often limit usage based on tokens per minute, with separate limits for input and output.
That means an agent can hit rate limits even while making relatively few requests if each request contains a large conversation history.
So context size directly affects rate limiting.
Parallel subagents make this worse because each may send its own large context.
A production harness therefore needs:
- concurrency limits,
- shared token budgets,
- context management,
not merely retry logic after receiving a 429.
6.3 Streaming can fail halfway through a response
A streamed request may successfully deliver part of an answer and then lose its connection.
The core rule is:
Partial output is not output. This follows the same at-least-once trade-off described in message delivery and processing: retry the complete operation rather than treating an incomplete result as success.
If a stream breaks halfway through, discard the incomplete response entirely.
Do not append partial text to conversation history, and definitely do not execute a partially received tool call.
Instead, retry the complete model request.
Because repeated prompts can often benefit from provider-side caching, retrying the whole request may also be much cheaper than it initially appears.
6.4 Long model requests need long timeouts
Frontier models working on large contexts may legitimately take minutes to complete a request.
Timeout settings designed for ordinary REST APIs may therefore kill perfectly healthy model requests.
Production harnesses should use generous timeout limits and rely on streaming activity to determine whether the model is still alive.
It is also important to log whether a request failed because:
- the model/provider failed,
- or the client’s own timeout/socket configuration terminated it.
Otherwise, engineers may mistakenly diagnose networking problems as model behavior problems.
7. The wrapper
The appendix argues that retry handling does not require redesigning the entire agent harness.
Instead, the network layer can be wrapped with two functions:
call_once
This function performs one API request.
It converts the outcome into either:
- a successful JSON response,
- or an
ApiError.
That error records whether the problem was:
- an HTTP error,
- or a network failure.
It also determines whether the failure is retryable.
call_with_retries
This function repeatedly calls call_once.
If the error is retryable, it waits and tries again.
If the error is permanent—or the maximum number of attempts has been reached—it immediately raises the real error.
The delay uses either the provider’s retry-after value or exponential backoff with randomness.
The rest of the harness barely changes: the original network request is simply replaced with the retrying wrapper.
8. Two easy implementation mistakes
The first common mistake is defining which errors are retryable incorrectly.
Network failures and known temporary status codes should normally retry.
Errors such as 400 or 401 should fail immediately because repeated requests will not fix a malformed request or bad credentials.
The second mistake is catching too narrow a set of networking exceptions.
A connection can fail through exceptions such as http.client.HTTPException, including RemoteDisconnected.
Only catching urllib.error.URLError can therefore miss exactly the type of dropped connection the retry wrapper was supposed to handle.
9. Policies outside the retry function
Two important rules cannot be solved purely inside a retry helper.
First, the system needs a token-aware concurrency limit above its subagent or parallel execution layer.
Second:
Discard broken partial streams and retry the entire request.
SDKs and API gateways may automate much of the retry plumbing, but these higher-level architectural policies still belong to the harness designer.