# What Is an Agent Harness? (And Why Every Harness Needs a Control Plane)

The model is the smallest part of your agent. Everything around it — the loop, the tools, the memory, the budget — is the harness, and it's where reliability, cost, and risk actually live. Here's what an agent harness is, why harness engineering became a discipline, and the one thing every harness in production still needs.

**An agent harness is everything around the model that turns it into a working agent: the loop that runs it, the tools it can call, the context and memory fed into each turn, the verification, the budgets and limits, and the logging. The model decides; the harness acts — and the harness is where cost, reliability, and risk actually live.**

There's a line going around that names the thing everyone building agents has been circling: **Agent = Model + Harness.**

The model is the part everyone talks about. The harness is everything else, and it turns out to be most of the system.

<figure style="margin:30px 0;" role="img" aria-label="An agent equals a small model plus a large harness. The model decides; the harness — loop, tools, context and memory, verification, budgets and limits, and logging — does everything else, and is most of the system.">
<svg viewBox="0 0 680 340" width="100%" style="max-width:680px;height:auto;font-family:ui-sans-serif,system-ui,sans-serif;">
  <text x="0" y="20" fill="var(--color-text-primary)" font-size="15" font-weight="700">An agent is mostly harness</text>

  <text x="0" y="190" fill="var(--color-text-secondary)" font-size="13" font-weight="600">Agent =</text>

  <!-- Model: the smallest part -->
  <rect x="66" y="146" width="96" height="80" rx="9" fill="var(--color-accent)"/>
  <text x="114" y="181" fill="#0b0d10" font-size="14" font-weight="700" text-anchor="middle">Model</text>
  <text x="114" y="200" fill="#0b0d10" font-size="10.5" text-anchor="middle" opacity="0.85">decides</text>

  <text x="182" y="193" fill="var(--color-text-secondary)" font-size="20" font-weight="600" text-anchor="middle">+</text>

  <!-- Harness: everything else -->
  <rect x="204" y="52" width="468" height="248" rx="12" fill="none" stroke="var(--acp-border-strong)" stroke-width="1.5"/>
  <text x="222" y="80" fill="var(--color-text-primary)" font-size="12" font-weight="700" letter-spacing="0.04em">HARNESS <tspan fill="var(--color-text-muted)" font-weight="400">— acts, and it's most of the system</tspan></text>

  <g font-size="12" text-anchor="middle">
    <rect x="222" y="96" width="136" height="86" rx="8" fill="var(--color-surface)" stroke="var(--color-border)"/>
    <text x="290" y="144" fill="var(--color-text-secondary)">Loop</text>
    <rect x="370" y="96" width="136" height="86" rx="8" fill="var(--color-surface)" stroke="var(--color-border)"/>
    <text x="438" y="144" fill="var(--color-text-secondary)">Tools</text>
    <rect x="518" y="96" width="136" height="86" rx="8" fill="var(--color-surface)" stroke="var(--color-border)"/>
    <text x="586" y="138" fill="var(--color-text-secondary)">Context</text>
    <text x="586" y="154" fill="var(--color-text-secondary)">&amp; memory</text>
    <rect x="222" y="192" width="136" height="86" rx="8" fill="var(--color-surface)" stroke="var(--color-border)"/>
    <text x="290" y="240" fill="var(--color-text-secondary)">Verification</text>
    <rect x="370" y="192" width="136" height="86" rx="8" fill="var(--color-surface)" stroke="var(--color-border)"/>
    <text x="438" y="234" fill="var(--color-text-secondary)">Budgets</text>
    <text x="438" y="250" fill="var(--color-text-secondary)">&amp; limits</text>
    <rect x="518" y="192" width="136" height="86" rx="8" fill="var(--color-surface)" stroke="var(--color-border)"/>
    <text x="586" y="240" fill="var(--color-text-secondary)">Logging</text>
  </g>

  <text x="0" y="328" fill="var(--color-text-muted)" font-size="11.5">The model decides what to do. The harness does it — and holds the cost, the risk, and the surprises.</text>
</svg>
</figure>

## The model is the smallest part

Give a frontier model a prompt and it returns text. That's not an agent. An agent does things: it reads a ticket, searches the web, queries a database, drafts a reply, sends it — looping, checking its own work, deciding what to do next.

None of that is the model. The model decides; the **harness** acts. The harness is:

- the **loop** that runs the model again and again until the task is done,
- the **tools** it can call, and how their results are fed back in,
- the **context and memory** — what gets loaded into each turn, and what gets trimmed,
- the **verification** — the checks that run before a result reaches a user,
- the **budgets and limits** — how much a run may spend, how long it may go,
- and the **logging** of everything that happened.

Swap the model for a better one and a bad harness still produces a bad agent. Keep the model and fix the harness, and the same model ships. Practitioners have a phrase for it: the gap between what a model *can* do and what you actually *see* it do is a **harness gap**. Closing that gap is **harness engineering** — its own discipline now, with [Thoughtworks](https://martinfowler.com/articles/harness-engineering.html), [Databricks](https://www.databricks.com/blog/ai-harness), [LangChain](https://www.langchain.com/blog/the-anatomy-of-an-agent-harness), and [MongoDB](https://www.mongodb.com/company/blog/technical/agent-harness-why-llm-is-smallest-part-of-your-agent-system) all writing about it.

## Harness engineering is where the cost and the surprises live

Once you see agents as harnesses, the weird behavior makes sense.

An agent that costs 1.2¢ one run and 2.2¢ the next isn't a flaky model — it's a harness whose loop ran a different number of turns. The bill that balloons on a "simple" task is the harness **re-reading its whole context on every turn**; on most agents, the majority of the token spend is re-reading, not new work. The run that takes three different paths to the same goal is the harness's loop diverging. These are harness properties, and they're measurable: turns, re-read multiple, path variance, cost per step.

Harness engineering turns an agent's behavior into something you can see and tune instead of a black box that hands back an answer.

One harness component deserves its own pattern language: the control layer — the mechanisms that constrain what the agent *may* do rather than what it *can* do. We have catalogued the eight patterns every harness converges on (prompt rules, permission prompts, pattern rules, approval classifiers, hooks, sandboxes, proxies, and independent ledgers) in [harness control patterns](/harness-control-patterns).

## But the harness creates problems it can't solve itself

The harness diagrams leave a part out. The moment your harness is calling real tools, spending real money, and acting on behalf of real users, you inherit a set of problems that live below the harness — at the boundary where it touches your tools and your backends:

- **Identity.** When the harness calls your API, who is it acting as? The user who triggered it? The agent? The framework's service account? (This is the [three-party problem](/blog/the-three-party-identity-gap) — and most harnesses paper over it with one shared key.)
- **Authority.** What is this harness allowed to do? Can it delete, pay, email the whole company? A harness will call any tool you hand it; it has no opinion about which calls are dangerous.
- **Cost.** What is it spending — per agent, per tool call — and can you cap it before a runaway loop bills you $400?
- **Control.** Can you stop it? Block one caller who's abusing it? Deny one tool without a redeploy?
- **Evidence.** Can you show an auditor every action it took, as a specific user, with the result?

These aren't harness engineering problems. You can build a good harness and still have all five. They're **governance** problems, and they're cross-cutting — they apply to every tool call the harness makes, no matter which framework it's built in.

The field is starting to draw the line explicitly: observability tells you what the harness did; governance controls what it's allowed to do. Traces and logs are read-only. Identity, policy, budgets, and blocks are enforcement. A harness needs both, and harness engineering by itself gives you neither.

## Every harness needs a control plane

This is why the harness conversation keeps arriving at the same place — [the AI control plane](https://medium.com/@adnanmasood/agent-harness-engineering-the-rise-of-the-ai-control-plane-938ead884b1d).

A **control plane** sits between the harness and the tools it calls, and handles exactly the five things the harness can't:

- it gives the agent a **verified identity** on every call,
- it enforces **policy** — allow, deny, or require approval, per tool and per action,
- it **prices** every call, so you see the loop tax and can cap it,
- it gives you a **kill switch** — for the whole agent, or for one caller,
- and it writes the **audit trail** — every action, attributed, with its result.

Because the control plane lives at the tool-call boundary, it works across any harness: Claude Code, LangGraph, CrewAI, the Anthropic SDK, or one you wrote yourself.

Anyone shipping a real agent ends up engineering a harness. A harness in production also needs a control plane under it. Call it a control plane, a governance layer, an agent gateway — the name matters less than the layer existing. You engineer the harness for capability; you put a control plane under it for identity, cost, and control.

*(Full disclosure: we build one — the [Agentic Control Plane](/) — which is why we've spent so long staring at harnesses. But the argument stands whatever you run.)*
