An agent harness is the software around a model that turns it into an agent: it defines the tools the model may call, executes those calls, decides what the model sees next, enforces limits on what it may do and spend, and records every step. The model generates text. The harness does everything else, including every part that can damage something.
The distinction gets lost because products are named after models. When a team asks whether an agent is safe to run against its systems, it is asking about the harness, and a model card cannot answer it.
What an agent harness contains
Implementations differ in shape but converge on the same components.
- Tool definitions. A schema for each action the model may request: name, parameters, what it returns. The model never executes anything; it emits a request that matches a schema.
- Tool execution and sandboxing. The code that validates a request, runs it and returns the result. Where it runs, as which user, with which filesystem and network: that is the sandbox.
- Permissions. Which tools, on which resources, with which approvals. Read-only versus write, scoped paths, allowlisted commands, no credentials in the model's view.
- Context management. What goes into the model's window on each step and what is dropped, truncated or summarised. This is where agents lose the thread.
- Retries and budgets. How tool failures and model failures are retried, and ceilings on steps, tokens, wall-clock time and money.
- Logging and replay. A trace of every prompt, tool call, result and decision, stored so the run can be inspected and replayed.
- Evaluation hooks. Checks that run on intermediate and final outputs: tests, linters, policy rules, stop conditions.
The SWE-agent paper calls the tool-facing part of this an agent-computer interface and shows that its design, with the model held constant, changes what the agent can complete. That is the harness thesis in one experiment.
Tools and permissions: the capability boundary
A tool definition is a contract. The model produces a structured call; the harness checks it against the schema, checks it against policy, executes it and returns the output. The Model Context Protocol is one widely used convention for describing tools so that the same definitions work across harnesses.
Permissions are enforced at execution, not by instructing the model. A system prompt that says "do not delete files" is a request; a sandbox with no delete permission is a boundary. The two are not interchangeable, and a security review should be able to tell which one a product relies on.
Tool output is untrusted input. A file the agent reads or a page it fetches can contain text that reads like instructions, and the model cannot reliably tell the difference. The OWASP Top 10 for LLM applications lists this as prompt injection. The harness-level defences are scoping what tools can reach and requiring approval for consequential actions, not better wording in the prompt.
Context, retries and budgets
The context window is finite and every tool result competes for it. The harness decides whether a long file is shown whole, truncated or summarised, and which earlier steps are kept when the window fills. A harness that truncates test output at the wrong point produces an agent that keeps fixing the wrong failure.
Retries need two separate policies. A tool failure, such as a flaky test or a transient network error, can be retried immediately with a cap. A model failure, where the step produced nothing usable, needs a budget rather than a count, for the same reasons that LLM traffic needs a retry budget: each retry costs tokens and time precisely when things are already going badly.
Budgets are the stop conditions a model cannot set for itself: maximum steps per task, maximum tokens, maximum wall-clock time, maximum spend. An agent without budgets does not fail; it loops. The harness ends the run, records why, and hands back whatever partial result exists.
Logging, replay and evaluation hooks
Every step should be recorded: the prompt as sent, the tool call as emitted, the result as returned, and the decision that followed. The fields worth keeping and the ones to redact are the same as for any model traffic, covered in what to log for LLM traffic, with tool calls and their results added.
Replay is what makes the log useful. A recorded trace can be stepped through, with the recorded tool results, to find where a run went wrong, compared against another run of the same task, or turned into a regression test for the harness itself.
Evaluation hooks are the harness's own checks on the model's work: run the tests before accepting a diff, validate structured output against a schema, apply a policy rule before an external action. They are the difference between an agent that reports success and an agent whose success was checked.
Why the harness, not the model, decides whether an agent is safe
Take one model and two harnesses. The first gives it an unrestricted shell on a developer's machine with their credentials in the environment, no budget and no log. The second gives it scoped tools in a container, a spend ceiling, a trace, and a rule that its output lands as a diff for a person to approve. Same model, same quality of reasoning, completely different risk.
The questions a company asks before letting an agent near its systems are all harness questions: what can it touch, what can it spend, what did it do, can we stop it, and can we prove all of that afterwards. Model choice moves quality. Harness design moves blast radius.
Where a gateway fits
The harness makes model calls, and a gateway can sit in front of them. An LLM gateway handles key custody, per-key quotas, failover between providers and cost attribution for those calls. It does not execute tools, manage context or see the harness's decisions; it sees requests and responses.
That makes the gateway optional from the harness's point of view and useful from the organisation's. A team that already holds provider keys centrally and attributes model spend per key can route an agent's calls through the same control plane. That is how Runix Router is positioned in front of the agents on the Runix stack: optional, with the agent running independently of it.
Two agents exist on the agents layer today, both in development: Runix Code for software engineering and Runix Comic for comic dramas. Each is designed to run as an independent agent for its domain with its own harness, so the questions above, about tools, permissions, budgets and review gates, are answered per agent rather than by a shared runtime.
Questions this raises
What is an agent harness?
The runtime around a model that turns it into an agent: tool definitions and execution, permissions and sandboxing, context management, retries and budgets, logging and replay, and evaluation hooks. The model generates; the harness acts and enforces.
Is an agent harness the same as an agent framework?
They overlap. A framework is a library for building the loop; the harness is the deployed runtime with the specific tools, permissions, budgets and logs that one agent runs under. The security properties live in the harness.
Does an agent need an LLM gateway?
No. A gateway sits in front of the harness's model calls for key custody, quotas, failover and cost attribution; it does not execute tools or manage context. It is useful when an organisation already governs model traffic that way.