Runtime as featured inForbesRead the article

engineering

How to Build an AI Agent Harness with the Runtime API

A step-by-step tutorial for building your own AI agent harness in Python and TypeScript, with the Runtime API as the backend for sandboxes, coding agents, guardrails, approvals and evals.

Carlos Volante
Co-founder and CTO, Runtime
October 8, 2026 · 24 min read

An AI agent harness is the software around a model that turns it into an agent that can do real work. The model decides what to do next. The harness gives it a computer to work on, tools, instructions, guardrails, approvals, evals, and a record of every step.

Runtime is the AI agent harness for payment and fintech teams. This post shows how to build one: your own orchestration and product on top, with Runtime as the backend for sandboxes, coding agents, guardrails and evals.

You keep the parts that make your product yours: the UI, the routing, the business rules, the queue. Runtime runs the parts that are slow and risky to build: isolated sandboxes, coding agents, policy enforcement, approvals, grading, and fallbacks when a provider goes down. All code below calls the Runtime Cloud API directly from Python and TypeScript. IDs are illustrative.

What an agent harness does

A model on its own is a function. Text goes in, a decision comes out: answer, call this tool, run this command. The harness is everything that turns that decision into work, in a loop:

  1. Trigger. Something starts a run: a ticket, a Slack message, a schedule, an API call from your app.
  2. Context. The harness loads instructions, skills, tools and credentials, and hands the model the task.
  3. Act in a sandbox. The model picks an action and the harness executes it on an isolated computer.
  4. Check. Every action is checked against policy before it runs.
  5. Approve. Steps with consequences wait for a human.
  6. Record. The harness stores what happened, grades the result, and reports back.

The loop itself is a few hundred lines. The hard part is what surrounds step 3. For the non-technical version of this idea, see our guide on what an agent harness is. This post is the build.

The architecture

Text
+-----------------------------------------------------------+
|  Your app (you own this)                                  |
|  UI, routing, business logic, queue, auth, audit export   |
+-----------------------------+-----------------------------+
                              |  HTTPS (REST + SSE)
+-----------------------------v-----------------------------+
|  Runtime API                                              |
|  sessions, templates, agents, prompts, events, files,     |
|  guardrails, approvals, grades, scorecards, telemetry     |
+-----------------------------+-----------------------------+
                              |
+-----------------------------v-----------------------------+
|  Runtime runs this                                        |
|  sandboxes across providers, coding agents, models        |
+-----------------------------------------------------------+
You ownRuntime runs
When a run starts and with what inputAn isolated VM per session
Your UI and who sees whatCoding agents (Claude Code, Codex, OpenCode, others)
Business rules and routing between agentsGuardrails enforced at the VM level
What happens with the resultApproval records, grades, scorecards, audit
Your queue, retries, and SLAsFallbacks across sandbox and model providers

Step 1: Authenticate

Create an organization API key in the dashboard under Settings > API Keys and store it in an environment variable. Keys start with runtm_sk_. Every request sends it as a bearer token, and org-scoped calls also send X-Organization-Id.

Errors return a JSON body with a detail field. A 429 carries a Retry-After header, and the docs recommend backing off on retries. Here is a small client that handles both, plus 5xx. The cURL tab verifies the key.

import os
import time
import requests
 
BASE = "https://app.runtm.com/api/cloud"
HEADERS = {
    "Authorization": f"Bearer {os.environ['RUNTM_API_KEY']}",
    "X-Organization-Id": os.environ["RUNTM_ORG_ID"],
}
 
def runtm(method, path, retries=3, **kwargs):
    headers = {**HEADERS, **kwargs.pop("headers", {})}
    for attempt in range(retries + 1):
        resp = requests.request(method, f"{BASE}{path}", headers=headers, timeout=60, **kwargs)
        if resp.status_code != 429 and resp.status_code < 500:
            break
        if attempt < retries:
            wait = int(resp.headers.get("Retry-After", 2 ** attempt))
            time.sleep(min(wait, 60))
    if resp.status_code >= 400:
        try:
            detail = resp.json().get("detail")
        except ValueError:
            detail = resp.text
        raise RuntimeError(f"Runtime {resp.status_code}: {detail}")
    return resp

Give each integration its own key with the smallest set of scopes it needs. A key can never exceed the role of the person who created it.

Step 2: Define the environment

A template is the environment every run boots from: repos, tooling, services, instructions, skills, and the secret names it expects. Values resolve from team secrets at boot, so they never sit in the template. This is one-time setup, usually done by an admin.

Create the template from the repo that holds your runbooks and scripts:

template = runtm("POST", "/org-templates", json={
    "name": "payout-investigation",
    "display_name": "Payout Investigation",
    "github_repo": "acme/payment-runbooks",
    "github_branch": "main",
    "agents": ["claude-code", "codex"],
    "tier": "standard",
}).json()
TEMPLATE_ID = template["id"]  # e.g. tmpl_abc123

Add guardrails scoped to that template. Allowlist rules take allow, ask or deny, and deny wins. Network rules limit what the sandbox can reach. Both are enforced at the VM level, so the agent can't talk its way around them.

runtm("POST", f"/org-templates/{TEMPLATE_ID}/guardrails", json={
    "type": "allowlist_rule_v0",
    "name": "no-direct-payout-release",
    "content": {"kind": "deny", "pattern": "payouts release*", "purpose": "Releases go through an approved script"},
})

Your SOPs become skills: a SKILL.md with a name, when to use it, and the steps. Attach skills before you build, because sessions boot from the snapshot. Attaching a skill to a template is documented in the CLI, so this step uses it:

Terminal
SKILL_ID=$(runtm-api skills create --name payout-investigation --md ./SKILL.md | jq -r .directive.id)
runtm-api skills attach "$SKILL_ID" --template tmpl_abc123

Then build the snapshot once:

build = runtm("POST", f"/org-templates/{TEMPLATE_ID}/build").json()
print(build["status"], build["build_id"])  # building build_xyz789

Optionally, give the work an identity. A roster agent carries instructions, a default template and harness, a rubric and a budget. Sessions launched as that agent inherit all of it.

agent = runtm("POST", "/v1/agents", json={
    "name": "Payout Investigator",
    "system_instructions": "Investigate held payouts. Write findings to /home/user/out/<payout_id>.json.",
    "default_template": "payout-investigation",
    "default_agent": "claude-code",
}).json()
AGENT_ID = agent["id"]

Step 3: Start a session

A session is one isolated VM with a filesystem, a terminal and a coding agent. Create one per unit of work. Pass an Idempotency-Key so a retry after a network error doesn't create a second sandbox.

session = runtm("POST", "/sessions", json={
    "agent": "claude-code",
    "template_id": TEMPLATE_ID,
    "name": "Payout po_7731",
    "visibility": "team",
}, headers={"Idempotency-Key": "payout-po_7731"}).json()
SESSION_ID = session["id"]

The agent field picks the coding agent. claude-code, codex and opencode run through the prompt API, which is what a headless harness wants. GitHub Copilot, Cursor, Devin and Gemini run interactively in the session terminal. You can change harness per session without touching the template's skills.

A new session takes a few seconds to provision. Poll until it's ready:

def wait_ready(session_id, timeout=120):
    deadline = time.time() + timeout
    while time.time() < deadline:
        s = runtm("GET", f"/sessions/{session_id}").json()
        if s["state"] in ("active", "running"):
            return s
        if s["state"] in ("error", "destroyed"):
            raise RuntimeError(f"Session {session_id} is {s['state']}")
        time.sleep(2)
    raise TimeoutError(session_id)

Step 4: Send a prompt and stream events

POST /sessions/{id}/prompt returns a text/event-stream. Each event has a type (text, tool_use, tool_result, error, done) and content. Only one prompt runs per session. If one is already running you get 202 with status: "already_running".

type PromptEvent = { type: "text" | "tool_use" | "tool_result" | "error" | "done"; content: string };
 
export async function* streamPrompt(sessionId: string, prompt: string): AsyncGenerator<PromptEvent> {
  const res = await runtm(`/sessions/${sessionId}/prompt`, {
    method: "POST",
    headers: { Accept: "text/event-stream" },
    body: JSON.stringify({ prompt }),
  });
  if (res.status === 202) throw new Error("A prompt is already running in this session");
 
  const reader = res.body!.pipeThrough(new TextDecoderStream()).getReader();
  let buffer = "";
  while (true) {
    const { value, done } = await reader.read();
    if (done) return;
    buffer += value;
    const lines = buffer.split("\n");
    buffer = lines.pop() ?? "";
    for (const line of lines) {
      if (line.startsWith("data:")) yield JSON.parse(line.slice(5));
    }
  }
}
 
for await (const event of streamPrompt(sessionId, "Why is payout po_7731 still on hold? Check the ledger and the processor.")) {
  if (event.type === "text") process.stdout.write(event.content);
  if (event.type === "tool_use") console.log(`\n[tool] ${event.content}`);
  if (event.type === "error") throw new Error(event.content);
  if (event.type === "done") break;
}

For a UI that watches a session from another process, subscribe to GET /sessions/{id}/events. It streams prompt_progress, file_changed, process_output, state_changed and heartbeat events.

import json
import httpx
 
def watch(session_id):
    url = f"{BASE}/sessions/{session_id}/events"
    with httpx.stream("GET", url, headers={**HEADERS, "Accept": "text/event-stream"}, timeout=None) as r:
        event = None
        for line in r.iter_lines():
            if line.startswith("event:"):
                event = line[6:].strip()
            elif line.startswith("data:"):
                data = json.loads(line[5:])
                if event == "file_changed":
                    print("file", data["action"], data["path"])
                elif event == "state_changed":
                    print("state", data["previous_state"], "->", data["state"])
                elif event == "prompt_progress":
                    print(data.get("content", ""), end="", flush=True)

If you don't need live output, submit and poll instead. The session object carries last_prompt.status, which ends as completed, error or timed_out.

runtm("POST", f"/sessions/{SESSION_ID}/prompt", json={"prompt": "Investigate payout po_7731"})
while (status := runtm("GET", f"/sessions/{SESSION_ID}").json().get("last_prompt", {}).get("status")) == "running":
    time.sleep(3)

Step 5: Read results and files

Prose is hard to parse. Have the skill tell the agent to write its result to a known path as JSON, then read the file. Your app gets structured output, and the file stays in the session for anyone reviewing the run.

listing = runtm("GET", f"/sessions/{SESSION_ID}/files",
                params={"path": "/home/user/out", "recursive": True}).json()
for f in listing["files"]:
    print(f["type"], f["path"], f["size"])
 
report = runtm("GET", f"/sessions/{SESSION_ID}/file",
               params={"path": "/home/user/out/po_7731.json"}).json()
finding = json.loads(report["content"])

The full transcript is available too. Use it to rebuild a view after a disconnect or to show reviewers what the agent did.

history = runtm("GET", f"/sessions/{SESSION_ID}/history").json()
print(f"{history['count']} events in the transcript")

Step 6: Human approvals

An approval is a gate the agent raises from inside its own run, right before a step with consequences. There is no public endpoint that creates one. The agent calls the runtm-approval helper from its runbook:

Terminal
# In SKILL.md, immediately before the step that sends the reply
runtm-approval request --kind customer_reply \
  --message "Draft reply for payout po_7731 is at out/po_7731.md. Approve to send it." \
  --required-role admin --wait

Exit code 0 means approved, 1 rejected, 4 timed out. On rejection the runbook stops and reports the note. Pair it with a deny rule on the risky command, as in Step 2, so the approval is a control and not a suggestion.

Your app surfaces pending approvals wherever your reviewers work, and resolves them over the API:

pending = [a for a in runtm("GET", f"/sessions/{SESSION_ID}/approvals").json()
           if a["status"] == "pending"]
 
for a in pending:
    show_in_review_queue(a["id"], a["kind"], a["message"])  # your UI
 
def resolve(approval_id, approve, note):
    return runtm("POST", f"/sessions/{SESSION_ID}/approvals/{approval_id}/resolve",
                 json={"action": "approve" if approve else "reject", "note": note}).json()

Resolution is role-gated on the server: admins and owners always can, otherwise the resolver must match required_role or required_team_id. That check runs against the key making the call, so if your service resolves on behalf of people, enforce in your app who may click Approve. Runtime doesn't send approvals to Slack. Approvers use your UI, the session card in the dashboard, or the CLI (runtm-api session approvals resolve).

Step 7: Webhooks for long-running work

Some runs take minutes. Instead of holding a stream open, configure an outbound webhook URL in the dashboard preferences. Runtime emits two events, prompt.completed and prompt.timed_out:

JSON
{
  "event": "prompt.completed",
  "session_id": "9f3a3f22-1d4e-4a9a-9a1f-3e5c6b1a0c11",
  "timestamp": "2026-10-08T15:18:54Z",
  "data": { "prompt_preview": "Investigate payout po_7731", "model": "sonnet", "cost_usd": 0.07, "duration_seconds": 23, "status": "completed" }
}

Delivery has a 5 second timeout and no retries. The docs don't describe a signature header, so treat the webhook as a hint: acknowledge fast, then confirm by reading the session from the API.

TypeScript
import express from "express";
const app = express();
app.use(express.json());
 
app.post("/webhooks/runtime", async (req, res) => {
  res.sendStatus(200);
  const { event, session_id } = req.body;
  const session = await runtm(`/sessions/${session_id}`).then((r) => r.json());
  if (session.last_prompt?.status === "completed") await finishJob(session_id);       // your code
  if (event === "prompt.timed_out" || session.last_prompt?.status === "timed_out") await retryJob(session_id);
});

Because there are no retries, keep a slow poll of open jobs as a backstop.

Step 8: Hand off to another agent

Work often spans agents. A payout investigator may need a ledger check from a reconciliation agent with different permissions. Launch a session as that roster agent with agent_id, and the child boots with its instructions, template and harness:

child = requests.post("https://app.runtm.com/api/v0/sessions/launch", headers=HEADERS, timeout=60, json={
    "agent_id": "b277d18b-b2c8-4410-9e10-8c7b2a415d73",
    "prompt": "Reconcile yesterday's ledger and list the breaks",
})
child.raise_for_status()
print(child.json())

The call returns the child session ID immediately. Poll it until last_prompt.status is completed, error or timed_out, then read its history or output files as in Step 5. Your orchestrator can make this call, or an agent can make it from inside its own run.

Each agent controls who may launch it with a callable_by allowlist:

runtm("PATCH", "/v1/agents/b277d18b-b2c8-4410-9e10-8c7b2a415d73", json={
    "callable_by": {
        "humans": False,
        "agents": [{"agent_id": AGENT_ID, "approval": {"required_team_id": "team_ops"}}],
    },
})

With that edge, a handoff needs an approved approval from the ops group, passed as approval_id. Delegation is capped at three hops, and refusals come back as AGENT_NOT_CALLABLE, APPROVAL_REQUIRED or DELEGATION_TOO_DEEP.

Step 9: Measure and control cost

You need to know whether agents do the job, and whether they still do after a model update. Give the roster agent evaluation categories and a budget. Each category says when a run belongs, what a pass looks like, and what the task costs a person:

runtm("PATCH", f"/v1/agents/{AGENT_ID}", json={
    "evaluator_criteria": {"categories": [{
        "key": "held-payout",
        "description": "Someone asks why a payout is held",
        "success_criteria": "Names the hold reason, cites the ledger entry and processor status, proposes a next step",
        "human_cost_usd": 18,
        "human_minutes": 25,
    }]},
    "economics": {"budget": {"monthly_usd_cap": 400}},
})

Runs are graded after they complete. Read one verdict, the per-session cost, and the scorecard for every agent:

grade = runtm("GET", f"/sessions/{SESSION_ID}/grade").json()
if grade["graded"]:
    print(grade["success"], grade["reason"], "rubric v", grade["criteria_version"])
 
usage = runtm("GET", f"/sessions/{SESSION_ID}/usage").json()
print(usage["total_cost_usd"])
 
for a in runtm("GET", "/sessions/telemetry/agents", params={"days": 7}).json()["agents"]:
    print(a["name"], a["objective_hit_rate"], a["spend_usd"], a["budget"]["exceeded"])

The grader sees the final answer, not the tool calls, so ask the agent to include its evidence in the reply. The per-agent budget flags the agent on the scorecard and writes an audit event. It doesn't stop launches. Hard spend and concurrency limits are set at the organization level under Guardrails > Limits. Telemetry reads need the activity:read scope.

Lifecycle and reliability

Sessions that see no activity for 20 minutes pause automatically, and paused sessions cost no compute. Pause between bursts, resume when the next message arrives, and delete when the work is done:

runtm("POST", f"/sessions/{SESSION_ID}/pause")    # 409 if a prompt is running
runtm("POST", f"/sessions/{SESSION_ID}/resume")   # 410 if the sandbox expired
runtm("DELETE", f"/sessions/{SESSION_ID}")        # 204, permanent

Any mutating call, like a prompt or file write, also resumes a paused session. If a resume returns 410, create a fresh session. For long gaps the docs describe a heartbeat endpoint that resets the idle timer, now marked legacy in favor of a newer refresh endpoint. For unattended runs, set a TTL at creation so nothing lingers:

JSON
{ "agent": "claude-code", "template_id": "tmpl_abc123", "ttl_minutes": 60, "on_complete": "pause" }

Reliability goes deeper than retries. Runtime avoids depending on a single provider at any layer:

  • Sandbox providers. Sessions run on Google Kubernetes Engine (GKE), AWS Lambda microVMs, E2B, Daytona and Modal. If one has an issue, sessions run on another.
  • Models. Claude, OpenAI, Gemini and open-weight models, including ones served in your own cloud for PCI and PII work.
  • Coding agents. Claude Code, Codex, OpenCode and others. Switch an agent's harness without rewriting its skills.

This matters because of what the agents do. Runtime is live with customers that handle hundreds of billions of dollars in payments. A held payout touches support, payment ops, finance, and sometimes risk and compliance, and it often has a same-day deadline. If your harness depends on one sandbox vendor and one model API, an outage at either stops the investigation. With fallbacks at each layer, the work keeps going and the record stays complete.

Give your ops team a UI for free

Your harness can stay headless. Engineering owns the orchestration, the integrations and the guardrails. Everyone else doesn't have to wait for a frontend.

Runtime has a web app built for people who don't write code. Once engineering sets the templates, tools, credentials, approvals and roles, people in payment ops, risk, compliance and support build and run agents from their own SOPs inside those limits. The same agents are reachable from Slack, Teams, WhatsApp, email, SMS, voice, schedules and the API. Your app and the web app share one record of every run, with team-based access and RBAC across both.

Putting it together

Here is the whole loop in TypeScript and Python, using the runtm and streamPrompt (stream_prompt) helpers from above: create a session, stream a prompt, read the structured result, and clean up.

import { runtm, streamPrompt } from "./runtime";
 
const READY = new Set(["active", "running"]);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
 
export async function investigatePayout(payoutId: string) {
  const session = await runtm("/sessions", {
    method: "POST",
    headers: { "Idempotency-Key": `payout-${payoutId}` },
    body: JSON.stringify({ agent: "claude-code", template_id: "tmpl_abc123", name: `Payout ${payoutId}` }),
  }).then((r) => r.json());
 
  try {
    for (let i = 0; i < 60; i++) {
      const s = await runtm(`/sessions/${session.id}`).then((r) => r.json());
      if (READY.has(s.state)) break;
      if (s.state === "error" || s.state === "destroyed") throw new Error(`Session ${s.state}`);
      await sleep(2000);
    }
 
    const prompt = `Investigate payout ${payoutId}. Write your findings to /home/user/out/${payoutId}.json.`;
    for await (const event of streamPrompt(session.id, prompt)) {
      if (event.type === "error") throw new Error(event.content);
      if (event.type === "done") break;
    }
 
    const params = new URLSearchParams({ path: `/home/user/out/${payoutId}.json` });
    const file = await runtm(`/sessions/${session.id}/file?${params}`).then((r) => r.json());
    return JSON.parse(file.content);
  } finally {
    await runtm(`/sessions/${session.id}/pause`, { method: "POST" }).catch(() => {});
  }
}

We pause rather than delete so a reviewer can open the session and see exactly what happened. Delete it once your retention window passes.

Wrapping up

A harness is a product decision as much as an engineering one. Build the parts that are specific to your business: where work comes from, how it's routed, what your users see. Use Runtime for the parts that are the same for everyone and expensive to get wrong: isolated sandboxes, coding agents, guardrails, approvals, grading, and fallbacks across providers. The full API reference is at docs.runtm.com.

Build your harness on Runtime

Bring one workflow. We'll help you wire the Runtime API into your app, with guardrails, approvals and evals from the first run.