Runtime as featured inForbesRead the article

What Is an Agent Harness? How AI Agents Run in Production

An agent harness is the software around an AI model that gives it a computer, tools, memory, permissions, and an audit trail. How it works in payments.

Gus TrigosCo-founder and CEO, RuntimeUpdated October 4, 20266 min readconcepts
agent-harnessai-agentsagent-infrastructurepayment-operationshuman-in-the-loopaudit-trail

Every rock climber knows strength isn't what keeps you alive on a hard route. The harness does.

AI agents work the same way. The model is the strength. It can read a bank RFI, trace a payout across three systems, and draft a reply in seconds. But strength alone is how you fall. The thing that lets you climb higher, safely, is everything wrapped around the model.

That's the agent harness.

The short answer. An agent harness is the software that surrounds an AI model and turns it into an agent that can do real work. The model reasons. The harness gives it a computer to work on, tools and data to use, memory between runs, permissions on what it can touch, approvals before sensitive actions, and a record of every step. In production, the harness decides whether an agent is useful or dangerous. Runtime is an agent harness built for payment teams, where every agent runs on its own computer in your cloud, with your approvals and a full audit trail.

The Model Isn't the Agent

When people say "we built an agent," they usually mean they wrote a good prompt and connected an API. That works in a demo. It breaks the first time the agent hits a timeout, runs out of context, or tries to touch something it shouldn't.

The model only does one thing. It takes in text and decides what to do next. Everything else is harness work:

  • Where does the agent run, and what happens to that machine afterward?
  • Which systems can it read, and which can it change?
  • Who approves before money moves?
  • What happens when the model provider goes down mid-task?
  • How do you show an auditor exactly what the agent did last Tuesday?

None of those are model questions. All of them decide whether you can trust the agent on a real queue.

What's in a Harness

A production harness has a handful of parts. Skip one and you find out which one the hard way.

PartWhat it doesWhat breaks without it
Its own computerEach agent works in an isolated sandbox with a browser, terminal, and filesAgents share state, leak data between runs, or touch production directly
Tools and dataAPIs, databases, MCP servers, and a browser for portals with no APIThe agent can talk about the work but can't do it
Scoped credentialsMasked secrets the agent can use but never seeOne prompt injection exposes your keys
Access controlRole-based permissions for every agent and teammateAn agent built for support can read compliance files
ApprovalsSensitive steps pause until a person says yesMoney moves on a hallucination
Memory and skillsProcedures and lessons carried from one run to the nextEvery case starts from zero
Audit trailEvery query, tool call, approval, and cost, stored and exportableYou can't answer "why did the agent do that?"
EvalsTests against real resolved cases before any change shipsA model update quietly makes the agent worse
Routing and recoveryThe right model for each task, plus retry and recovery paths when infrastructure or provider calls failYou overpay for routine work and stall during outages

Each of these is a project on its own. Together, they're the difference between a demo and a teammate.

Why Payments Raises the Bar

Most agent harness writing assumes the worst case is a bad answer. In payments, the worst case is a bad action.

An agent that releases a held payout to the wrong account costs real money. An agent that misses a sponsor bank's RFI deadline puts the bank relationship at risk. An agent that copies a card number into a prompt creates a PCI problem. And when an examiner asks how a case was handled, "the AI did it" isn't an answer.

So the harness for a payment team has stricter jobs:

  • Run where the data lives. Agent computers in your own cloud, so PCI data never leaves your environment.
  • Keep card data out of the model. Masked credentials, plus filters that strip card numbers and SSNs from prompts and logs.
  • Start read-only. Agents investigate freely, and anything that moves money, applies a reserve, or replies to your bank waits for approval.
  • Keep the evidence. Every run stored end to end, ready when compliance, an auditor, or your sponsor bank asks.
  • Stay up. Build in retries, recovery paths, and provider-aware routing, because a transient outage shouldn't stop a payout investigation.

This is also why most internal builds stall. The first agent is easy. The approvals, audit trail, and access controls take a team of senior engineers a couple of quarters before that agent touches real data. Meanwhile the queue keeps growing.

How Runtime Works

Runtime is the agent harness for payment teams. Anyone on the team can build and run agents on it, and engineering keeps control of what each one can do.

  1. Describe the job. Someone on the team writes what the agent should do in plain English and attaches the SOP. Runtime builds the agent and gives it its own computer.
  2. Connect your tools. APIs, databases, and MCP servers, plus a browser for processor and bank portals with no API. Role-based access sets what each agent can read and do.
  3. Run it. Call the agent from Slack, Teams, email, text, or a phone call, or trigger it from an alert or a schedule. It investigates on its own computer and shows every step.
  4. Keep a human in the loop. Money movement, reserves, and replies to your bank wait for approval. Every run is stored with its evidence, cost, and outcome.

Runtime is neutral on models and agents. It runs Claude Code, Codex, and OpenCode with Claude, OpenAI, Gemini, or open-weight models served from your own cloud. Evals route routine work to the smallest model that still passes, so runs get cheaper over time. Over time, every run also becomes memory the next one can use. We call that a system of intelligence.

If you want to see it on a specific queue, the payment operations, compliance, and underwriting pages walk through real cases. The security page covers deployment and controls.

Frequently asked questions

What is an agent harness?

An agent harness is the software that surrounds an AI model and turns it into an agent that can do real work. The model reasons. The harness gives it a computer to work on, tools and data to use, memory between runs, permissions on what it can touch, approvals before sensitive actions, and a record of every step. Runtime is an agent harness built for payment teams.

What's the difference between an agent harness and an agent framework?

A framework is a library you use to write an agent's logic, like the loop that decides which tool to call next. A harness is what that agent runs inside in production: the isolated computer, credentials, access control, approvals, audit trail, evals, and failover. You can run agents built with many frameworks, or coding agents like Claude Code and Codex, inside the same harness.

If we already use Claude Code or Codex, do we still need a harness?

Claude Code and Codex are excellent agents, and they include their own local harness for a single developer. Running them for a team on production payment data adds requirements they don't cover on their own: isolated computers per run, scoped credentials, role-based access, approvals in Slack or Teams, a stored audit trail, and recovery paths when infrastructure or provider calls fail. Runtime runs Claude Code, Codex, and OpenCode inside that harness.

Can an agent harness run inside our own cloud?

It should, if the agent touches card data or PII. With Runtime, agent computers run in your own AWS, GCP, or Azure account while we manage the control plane, or you can self-host the whole platform. PCI data stays in your environment.

What does an agent harness cost compared to building one?

Building a production harness in-house usually takes a team of senior engineers a couple of quarters before the first agent touches real data, because approvals, audit trails, and access controls come first. A harness like Runtime is priced against the work it takes off your team, not per seat.

Clip In

Strong models are everywhere now. Anyone can rent the same strength.

What separates teams that get value from agents and teams that get burned is the harness around them. Pick the route your team actually needs to climb, the queue that keeps growing, and make sure you're tied in before you start.

Then climb.

Put an agent on your busiest queue

Bring one SOP. A forward-deployed AI engineer builds the first agent with your team, inside your cloud, with your approvals.