harnessengineering llm ai agents ai/agenticai

Core Idea

Model routing: assign each task to the cheapest model that does it reliably, and give weaker models more scaffolding.

  • Route by consequence, not difficulty. A yes/no safety classifier looks trivial but a wrong answer can let a dangerous command through, so it needs verification and fail-closed behavior.
  • Don’t route unverifiable work to the cheapest model: if the main agent trusts a wrong subagent result, fixing it costs more than the routing saved.
  • The weaker the model, the more structure the harness provides: checklists instead of open prompts, specialized tools (grep, find) instead of a shell, more steering.
  • Operational traps: provider parameter dialects, invisible calls (safety, compaction, recaps) that still cost tokens, and one pinned version per model job.

1. Introduction: Why use multiple models?

A production AI harness makes many different kinds of model calls during one session. Not every task requires the most powerful model.

For example, the main coding loop may require a frontier model, while simpler tasks such as generating status messages, summarizing conversation history, extracting information, or producing recaps can use cheaper models.

The core idea is model routing: assign each task to the cheapest model that can perform it reliably. This reduces both cost and latency without unnecessarily sacrificing quality.

2. The routing table

Different jobs should use different levels of models.

JobRecommended model level
Main reasoning/coding loopFrontier model
Search or verification.Mid-tier Model
Safety ClassificationSmall, fast model with verification
Conversation compaction summariesMid-tier model
Recaps, labels, status messagesSmallest suitable model
Extracting answers from web pagesSmall or mid-tier model

The table itself is less important than the two principles behind it.

Route by consequence, not just difficulty

A simple task can still be high-risk.

For example, a safety classifier may only answer “yes” or “no,” but a wrong decision could allow a dangerous command such as deleting files. Therefore, even though the task looks easy, it needs safeguards such as verification and fail-closed behavior.

By contrast, an incorrect status message or recap has little consequence.

So model selection should depend on how damaging a mistake would be, not simply how complicated the task appears.

Delegated tasks still need capable models

A common mistake is assigning subagent work to the cheapest possible model because it happens “in the background.”

If that model produces incorrect research, the main model may blindly trust it and build further reasoning on top of the mistake. Fixing the mistake later can consume more expensive frontier-model tokens than were saved.

The practical rule is:

  • Route down when results can be cheaply verified or mistakes are cosmetic.
  • Use stronger models when the main agent must trust the result without independently checking it.

3. Prompt tiers: the body adapts to the brain

Choosing the model is only half of routing. The harness should also change how it communicates with different model tiers.

A weaker model often needs more explicit structure and guidance than a strong frontier model.

System prompt tiers

The strongest models can work with relatively concise system prompts.

Mid-tier models may benefit from:

  • more examples,
  • clearer procedures,
  • stronger instructions.

Low-tier models may need instructions that resemble a step-by-step checklist.

The policies remain the same; only the amount of guidance changes.

Tool set tiers

Powerful models can often operate effectively through flexible tools such as a shell.

Smaller models perform better when given specialized tools like: grep,find,ls, as discussed in tool extension patterns.

These tools constrain the possible actions and make mistakes less likely.

Therefore, weaker models receive more structured tools.

Steering density

Small models also benefit from more frequent reminders.

Instructions that a frontier model might consider repetitive or distracting can be essential for keeping a weaker model on track.

The harness can therefore adjust how many reminders or steering messages it sends depending on the model tier.

General principle

There is a trade-off between model capability and scaffolding:

The weaker the model, the more structure the harness must provide.

A system supporting several model tiers therefore needs several versions of its prompting, tools, and guidance.

If a system uses only one frontier model, this complexity may not be necessary.

4. Operational traps

Using multiple models and providers introduces several engineering problems.

Dialect drift

Cheap models may come from different APIs, providers, or local inference systems.

Those providers may interpret parameters differently. A parameter accepted by one API may cause another API to fail.

Common examples include:

  • temperature settings,
  • reasoning controls,
  • optional request fields.

The system should explicitly handle each provider’s API differences and fail loudly rather than silently switching providers, especially for safety-sensitive tasks.

Account for invisible calls

Many model calls never appear in the visible conversation, including:

  • safety classification,
  • conversation summarization,
  • recaps,
  • internal processing.

They still consume tokens and money.
Therefore, usage tracking and cost accounting should include every model call, not just visible assistant responses.
Otherwise, cost dashboards will significantly underestimate actual usage.

Pin versions per job

Each model role should be treated as its own dependency.

For example, upgrading the main reasoning model should not automatically upgrade the safety classifier.

Each model used for a specific job should have:

  • its own pinned version,
  • its own testing,
  • its own upgrade process.

This prevents unrelated model changes from unexpectedly altering important behavior.

Let the user override

Model routing is ultimately a policy decision.

Users should be able to see and configure that policy rather than having model choices permanently hard-coded into the system.

This complexity is optional when cost and latency are unimportant. But once a system operates at scale, routing cheaper models to appropriate tasks can significantly reduce both expense and response time while preserving strong models for the tasks where they matter most. The same idea resembles mixture-of-experts routing, but the harness routes whole tasks while an MoE model routes tokens.