What Is an Agent Harness? How AI Agents Run in Production
An agent harness is the software around an AI model that gives it a computer, tools, memory, permissions, and an audit trail. How it works in payments.
Every rock climber knows strength isn't what keeps you alive on a hard route. The harness does.
AI agents work the same way. The model is the strength. It can read a bank RFI, trace a payout across three systems, and draft a reply in seconds. But strength alone is how you fall. The thing that lets you climb higher, safely, is everything wrapped around the model.
That's the agent harness.
The short answer. An agent harness is the software that surrounds an AI model and turns it into an agent that can do real work. The model reasons. The harness gives it a computer to work on, tools and data to use, memory between runs, permissions on what it can touch, approvals before sensitive actions, and a record of every step. In production, the harness decides whether an agent is useful or dangerous. Runtime is an agent harness built for payment teams, where every agent runs on its own computer in your cloud, with your approvals and a full audit trail.
The Model Isn't the Agent
When people say "we built an agent," they usually mean they wrote a good prompt and connected an API. That works in a demo. It breaks the first time the agent hits a timeout, runs out of context, or tries to touch something it shouldn't.
The model only does one thing. It takes in text and decides what to do next. Everything else is harness work:
- Where does the agent run, and what happens to that machine afterward?
- Which systems can it read, and which can it change?
- Who approves before money moves?
- What happens when the model provider goes down mid-task?
- How do you show an auditor exactly what the agent did last Tuesday?
None of those are model questions. All of them decide whether you can trust the agent on a real queue.
What's in a Harness
A production harness has a handful of parts. Skip one and you find out which one the hard way.
| Part | What it does | What breaks without it |
|---|---|---|
| Its own computer | Each agent works in an isolated sandbox with a browser, terminal, and files | Agents share state, leak data between runs, or touch production directly |
| Tools and data | APIs, databases, MCP servers, and a browser for portals with no API | The agent can talk about the work but can't do it |
| Scoped credentials | Masked secrets the agent can use but never see | One prompt injection exposes your keys |
| Access control | Role-based permissions for every agent and teammate | An agent built for support can read compliance files |
| Approvals | Sensitive steps pause until a person says yes | Money moves on a hallucination |
| Memory and skills | Procedures and lessons carried from one run to the next | Every case starts from zero |
| Audit trail | Every query, tool call, approval, and cost, stored and exportable | You can't answer "why did the agent do that?" |
| Evals | Tests against real resolved cases before any change ships | A model update quietly makes the agent worse |
| Routing and recovery | The right model for each task, plus retry and recovery paths when infrastructure or provider calls fail | You overpay for routine work and stall during outages |
Each of these is a project on its own. Together, they're the difference between a demo and a teammate.
Why Payments Raises the Bar
Most agent harness writing assumes the worst case is a bad answer. In payments, the worst case is a bad action.
An agent that releases a held payout to the wrong account costs real money. An agent that misses a sponsor bank's RFI deadline puts the bank relationship at risk. An agent that copies a card number into a prompt creates a PCI problem. And when an examiner asks how a case was handled, "the AI did it" isn't an answer.
So the harness for a payment team has stricter jobs:
- Run where the data lives. Agent computers in your own cloud, so PCI data never leaves your environment.
- Keep card data out of the model. Masked credentials, plus filters that strip card numbers and SSNs from prompts and logs.
- Start read-only. Agents investigate freely, and anything that moves money, applies a reserve, or replies to your bank waits for approval.
- Keep the evidence. Every run stored end to end, ready when compliance, an auditor, or your sponsor bank asks.
- Stay up. Build in retries, recovery paths, and provider-aware routing, because a transient outage shouldn't stop a payout investigation.
This is also why most internal builds stall. The first agent is easy. The approvals, audit trail, and access controls take a team of senior engineers a couple of quarters before that agent touches real data. Meanwhile the queue keeps growing.
How Runtime Works
Runtime is the agent harness for payment teams. Anyone on the team can build and run agents on it, and engineering keeps control of what each one can do.
- Describe the job. Someone on the team writes what the agent should do in plain English and attaches the SOP. Runtime builds the agent and gives it its own computer.
- Connect your tools. APIs, databases, and MCP servers, plus a browser for processor and bank portals with no API. Role-based access sets what each agent can read and do.
- Run it. Call the agent from Slack, Teams, email, text, or a phone call, or trigger it from an alert or a schedule. It investigates on its own computer and shows every step.
- Keep a human in the loop. Money movement, reserves, and replies to your bank wait for approval. Every run is stored with its evidence, cost, and outcome.
Runtime is neutral on models and agents. It runs Claude Code, Codex, and OpenCode with Claude, OpenAI, Gemini, or open-weight models served from your own cloud. Evals route routine work to the smallest model that still passes, so runs get cheaper over time. Over time, every run also becomes memory the next one can use. We call that a system of intelligence.
If you want to see it on a specific queue, the payment operations, compliance, and underwriting pages walk through real cases. The security page covers deployment and controls.
Frequently asked questions
What is an agent harness?
An agent harness is the software that surrounds an AI model and turns it into an agent that can do real work. The model reasons. The harness gives it a computer to work on, tools and data to use, memory between runs, permissions on what it can touch, approvals before sensitive actions, and a record of every step. Runtime is an agent harness built for payment teams.
What's the difference between an agent harness and an agent framework?
A framework is a library you use to write an agent's logic, like the loop that decides which tool to call next. A harness is what that agent runs inside in production: the isolated computer, credentials, access control, approvals, audit trail, evals, and failover. You can run agents built with many frameworks, or coding agents like Claude Code and Codex, inside the same harness.
If we already use Claude Code or Codex, do we still need a harness?
Claude Code and Codex are excellent agents, and they include their own local harness for a single developer. Running them for a team on production payment data adds requirements they don't cover on their own: isolated computers per run, scoped credentials, role-based access, approvals in Slack or Teams, a stored audit trail, and recovery paths when infrastructure or provider calls fail. Runtime runs Claude Code, Codex, and OpenCode inside that harness.
Can an agent harness run inside our own cloud?
It should, if the agent touches card data or PII. With Runtime, agent computers run in your own AWS, GCP, or Azure account while we manage the control plane, or you can self-host the whole platform. PCI data stays in your environment.
What does an agent harness cost compared to building one?
Building a production harness in-house usually takes a team of senior engineers a couple of quarters before the first agent touches real data, because approvals, audit trails, and access controls come first. A harness like Runtime is priced against the work it takes off your team, not per seat.
Clip In
Strong models are everywhere now. Anyone can rent the same strength.
What separates teams that get value from agents and teams that get burned is the harness around them. Pick the route your team actually needs to climb, the queue that keeps growing, and make sure you're tied in before you start.
Then climb.
Put an agent on your busiest queue
Bring one SOP. A forward-deployed AI engineer builds the first agent with your team, inside your cloud, with your approvals.