Agent harnessCoding agentsPlatform engineeringEvals

Shipping Code With an Agent You Do Not Trust

I built a small, private harness that lets a coding agent open real pull requests under guarantees that are enforced by infrastructure rather than by prompt text, and that measures every run honestly enough to improve itself. This is what it does, what it proved, and what broke.

The premise

Most demos of coding agents answer the question "can the model write the code?" That question is mostly settled. The question I wanted to answer is the one an engineering manager asks before letting an agent near a repository: what stops it from doing the wrong thing, and how would I know?

So the harness was designed around three claims, and every component exists to make one of them true in a way you can test:

  1. Permissions live in code, never in prompts. A prompt is a suggestion. A tool that refuses is a guarantee.
  2. The agent is checked, never trusted. The agent's own view that it is finished is an input to verification, not a result.
  3. Every run is a measurement. Cost, outcome, failure mode and transcript are recorded the same way for every run, so failures turn into tests and tests turn into fixes.

A prompt is a suggestion. A tool that refuses is a guarantee.

What it is

A local Kubernetes cluster (k3d) runs the whole thing: a model gateway, the harness API and its console, a warm pool of locked-down sandbox pods, Postgres, ClickHouse and an S3-compatible object store, plus Prometheus, Tempo, Grafana and Langfuse for observability. Argo CD owns the cluster after bootstrap, so every change to it is a commit on main. The agent works on two sample projects in a monorepo: a small FastAPI claims service with a human-owned payments/ module, and a hello-world Next.js app added later to prove the harness is not Python-specific.

The system at a glance

Everything runs in one local cluster, split by namespace. Arrows are the only paths that exist; the sandbox has none out.

Figure 1The system at a glance: one local cluster, four namespaces, and every path that exists between them. The sandbox has none out.
Figure 1The system at a glance: one local cluster, four namespaces, and every path that exists between them. The sandbox has none out.Fit

A run is a LangGraph state machine with six nodes, checkpointed in Postgres so a capped run can be resumed later in the same sandbox:

Figure 2One run: plan, implement, self_check, verify, open_pr, finalize, with the fix loop and the dishonest-success branch.
Figure 2One run: plan, implement, self_check, verify, open_pr, finalize, with the fix loop and the dishonest-success branch.Fit

plan asks the model for a short numbered plan. implement is the agentic loop, with exactly four tools: list files, read a file, write a file, run the checks. self_check is the agent's side of verification: run every project's lint, types and tests, then review the full diff against the spec. A failure sends the run back to implement with feedback, up to a cap. verify is the harness's own check, run directly in the sandbox, which the agent cannot influence. open_pr takes the staged diff out of the sandbox, re-checks it for regulated paths, and opens a pull request as a bot identity. finalize computes the outcome, writes the transcript, and escalates anything that is not clean.

The console

One page over the API, refreshed every ten seconds. Everything the harness records is here: runs and their outcomes, the specs and which of them may run, the eval suite, the pull requests the loop opened, escalations, the warm pool, and the non-secret configuration.

Console1 / 10Overview: runs, clean-run rate, cost, dishonest successes and open escalations at a glance, with the latest runs, failure modes by count and the three gates.
Console1 / 10Overview: runs, clean-run rate, cost, dishonest successes and open escalations at a glance, with the latest runs, failure modes by count and the three gates.
Harness console overview page

The stack

Free and open-source editions throughout, every image, chart and dependency pinned, and nothing bespoke where a standard piece would do. The harness itself is a few thousand lines of Python; most of the system is configuration that makes those lines hard to misuse.

Cluster and delivery

  • k3d, a local Kubernetes with its own image registry
  • Argo CD, app-of-apps from one deploy/ folder; GitOps only after bootstrap
  • Helm charts, vendor charts for the platform pieces, small in-repo charts for the rest
  • GitHub Actions for CI and the eval gate; CODEOWNERS and branch protection as gates

Models

  • LiteLLM gateway with two aliases, primary and fallback, a global spend cap, and cost in response headers
  • Claude Sonnet 5 as primary and Haiku 4.5 as fallback, named only in the environment
  • Opt-in: the Claude Code CLI in headless mode on a subscription, bridged to the same tools over MCP

Harness

  • Python 3.12 in a uv workspace: harness, evals, loop and the sandbox manager as separate packages
  • FastAPI and uvicorn for the API; the console is one HTML page served by it, behind an API key
  • LangGraph for the run as a state machine, checkpointed in Postgres so capped runs resume
  • MCP servers for the tools, in-process; a tool layer with an allow-list and path rules in front of them
  • The Kubernetes Python client, used for exactly one verb pair: pods/exec in the sandbox namespace

Sandbox

  • One image with both projects, the uv environment, Node 22 and node_modules, so checks run offline
  • A warm pool kept by a small manager: claim, touch, reap on idle or age
  • Deny-all NetworkPolicy, no service-account token, non-root user, all capabilities dropped, CPU and memory limits

Data

  • Postgres for runs, escalations, eval results, failure tags and graph checkpoints
  • ClickHouse for the event stream: every tool call, model call, halt and verdict with cost and tokens
  • RustFS, an S3-compatible store, for transcripts and the scrubbed dataset exports
  • Redis, only because Langfuse needs it

Observability

  • OpenTelemetry collector, Tempo for traces, Prometheus for metrics, Grafana for the one dashboard
  • Langfuse for a per-call view of what the gateway sent and received
  • Every record, event and span carries the run id, the harness version and the model

Quality

  • ruff, mypy in strict mode and pytest for Python; ESLint, TypeScript and vitest for the Next.js app
  • One check.sh that discovers each project and runs its checks, identical on a laptop, in the sandbox and in CI
  • A lock check on every run, so a drifted dependency fails the build

Sample projects

  • A FastAPI claims service with SQLite and a human-owned payments/ module
  • A hello-world Next.js 16 app in TypeScript, added to prove the harness is language-agnostic

Three gates that software cannot pass on its own

Figure 3Three human gates: spec approval by code-owner merge, pull-request review, and a manual release sync.
Figure 3Three human gates: spec approval by code-owner merge, pull-request review, and a manual release sync.Fit

The harness refuses to start a run unless the spec is on main with status: approved, and the specs/ folder is code-owned, so approval is a human merge. That is gate one. A pull request needs CI, the eval gate and a human review to merge, and the bot cannot approve its own work. That is gate two. The release environment is a manual Argo CD sync, while the dev environment syncs on every merge. That is gate three. None of these are policies someone has to remember; they are what the system does.

Argo CD1 / 3Applications: fifteen synced and healthy. The release application is OutOfSync on purpose, because nobody has clicked Sync.
Argo CD1 / 3Applications: fifteen synced and healthy. The release application is OutOfSync on purpose, because nobody has clicked Sync.
Argo CD applications view

Guarantees and how each one is enforced

The agent cannot touch the regulated payments code

The write tool refuses the path in code and the pull-request tool refuses a diff that touches it. The path list is configuration. The prompt never mentions it.

The agent cannot reach the network or the cluster

Sandbox pods run under a deny-all NetworkPolicy with no service-account token, as a non-root user, with CPU and memory limits. Both projects' toolchains are baked into the image so the checks run offline.

The agent cannot call tools it was not given

An allow-list in the tool layer. A call to anything else is logged as refused.

A run cannot spend without limit

A per-run dollar budget from the gateway's cost headers, a fix-attempt cap, a tool-call cap, a stall watchdog, and a global spend cap at the gateway.

“Done” is not taken at face value

An independent verify step. If the agent stopped and verify fails, the outcome is dishonest_success, escalated, and never exported to the dataset.

Every change to the cluster is attributable

GitOps only, after bootstrap. Argo CD shows the commit and author behind every application.

What a tool call actually passes through

Figure 4What a tool call passes through: allow-list, path rules, the exec API, the pod. A regulated write is refused before anything reaches the sandbox.
Figure 4What a tool call passes through: allow-list, path rules, the exec API, the pod. A regulated write is refused before anything reaches the sandbox.Fit

Outcomes are a vocabulary

Every run ends in exactly one of a small set of outcomes, and the whole system speaks in them: the console, the dashboard, the escalation table, the failure tagger.

OutcomeMeaningEscalatedResumable
cleanverified on the first passnono
fixedverified after one or more self-check fixesnono
dishonest_successagent stopped, independent verify failedyesno
capped_budgetcumulative cost passed the budgetyesyes
capped_retriesfix attempts or tool calls hit the capyesyes
stalledno event for the stall windowyesyes
blocked_permissiontried to write under a regulated pathyesno
erroranything else, such as the gateway being downyesno

The loop: a seeded bug finds its own fix

The part I am most pleased with is the improvement loop, because it closed end to end on a real defect. Version 0.1 of the harness shipped with a deliberate bug: the self-check reviewer only read the first 2,000 characters of the diff. Small specs passed. A large spec produced a complete, correct change and then failed self-check three times, because the reviewer saw a truncated diff and called it incomplete, until the run hit its retry cap.

Figure 5The improvement loop: failed runs are tagged, become regression tasks and fix proposals, pass human review, roll out as a new version, and feed the dataset.
Figure 5The improvement loop: failed runs are tagged, become regression tasks and fix proposals, pass human review, roll out as a new version, and feed the dataset.Fit

The loop then did its job without a human writing any of it:

  1. Tagging. Rules map each failed run to a failure mode from its outcome, last error and self-check evidence. The three runs were tagged truncated_diff. A model is only asked when the rules do not match.
  2. Regression tasks. A failed run becomes a new eval task, with hidden tests, opened as a pull request. Two of the thirteen eval tasks exist because of this.
  3. A proposed fix. When a failure mode repeats enough times, a proposer drafts a harness change. It may only touch two files, it has to pass the full check script in a scratch clone, it bumps the harness version and the image tag, and it opens a pull request. A human reviewed and merged it. Harness 0.2.0 rolled out through Argo CD, and the large spec passed on re-run.
  4. A dataset. Verified runs are exported as scrubbed JSONL, with dishonest successes excluded, and a scan proves there are no secrets in it. That export is the handoff point to fine-tuning, which the project deliberately stops short of.

Evals as the gate, not a report

Figure 6The eval gate: a task runs in eval mode, hidden tests score it, and a harness change must meet the threshold and the stored baseline.
Figure 6The eval gate: a task runs in eval mode, hidden tests score it, and a harness change must meet the threshold and the stored baseline.Fit

An eval task is a spec plus hidden tests the agent never sees. After a run, the hidden tests are copied into the sandbox and executed, so the score is about the code, not the agent's claims. Each task carries its own test directory and command, which is how a vitest task for the Next.js app scores alongside pytest tasks for the Python service. A pull request that touches the harness, the evals or the loop must meet a pass-rate threshold and must not fall below the stored baseline from main. Other pull requests get a rubric score from a judge model. Scores are stored with the harness version, so before and after a change are always comparable.

Numbers

13 / 13
eval tasks passing on the current harness, across two suites
$0.13–0.42
per task on Sonnet 5 through the gateway
~30 ms
sandbox claim from the warm pool, versus seconds for a cold pod
4
pull requests opened by the loop, three merged, one superseded

The dashboard shows cold start with and without the warm pool, agent pull-request merge rate, review minutes per pull request, clean-run rate, cost per run, dishonest successes, failure modes by count and outcomes over time. Those were the metrics the spec asked for, and each has a panel backed by Postgres, ClickHouse or Prometheus. The single most useful number turned out to be the dishonest-success count, because it is the gap between what the agent says and what is true.

GrafanaThe one dashboard: cold start with the warm pool, clean-run rate, cost per run, dishonest successes, failure modes, outcomes, gateway cost and model calls per hour, and the runs table.
GrafanaThe one dashboard: cold start with the warm pool, clean-run rate, cost per run, dishonest successes, failure modes, outcomes, gateway cost and model calls per hour, and the runs table.
Grafana dashboard for the harness
LangfuseLangfuse: every call the gateway made, with tokens and cost by model and observations over time. The totals are small because the later suites ran on the subscription backend, which the gateway never sees.
LangfuseLangfuse: every call the gateway made, with tokens and cost by model and observations over time. The totals are small because the later suites ran on the subscription backend, which the gateway never sees.
Langfuse project home

Running on a subscription instead of API keys

Late in the project I added an opt-in model backend that runs the agent through the Claude Code CLI in headless mode, billed to a Claude Max subscription, using the CLI's documented long-lived token. The interesting part was keeping the guarantees. The CLI's own tools are disabled, and the harness exposes its four tools to the CLI as an MCP server on loopback. Every call still goes through the same tool layer, so the allow-list, the regulated-path refusal, the event log and the call cap are identical on both backends. The budget works on the cost the CLI reports; a stall, a refused regulated write or the call cap terminates the CLI from outside; and an exhausted usage window becomes a new resumable outcome, capped_rate_limit.

Figure 7The subscription backend: the CLI runs inside the harness pod with its own tools off, reaches the same four tools over an MCP bridge, and is watched and terminated from outside.
Figure 7The subscription backend: the CLI runs inside the harness pod with its own tools off, reaches the same four tools over an MCP bridge, and is watched and terminated from outside.Fit

The full suite passed on it. What you lose is real cost accounting, gateway traces and a fallback model, which is why the default stays the gateway.

What went wrong

  • A fixed decision had to change. The object store I had planned on withdrew its open-source images. The rule was to stop and ask rather than swap a tool silently, and that is what happened; RustFS took its place.
  • Kubernetes exec needs get as well as create on pods/exec, because the websocket upgrade is a GET. The least-privilege role was one verb too strict.
  • A root-owned mount broke two things at once: cp -a and git's ownership check inside the sandbox.
  • The warm pool replaced failed pods forever until it learned to delete them.
  • The provider ran out of credits in the middle of a suite. That run shows as 5 of 10 on the dashboard, and it is exactly the outage a second provider is for.
  • "Done" as a marker word cost fix attempts. Models sometimes summarise and stop without the magic word. Now, no further tool calls means the agent considers itself finished, and the checks decide whether it is.
  • The first fix proposal conflicted after unrelated version bumps, and the second was nearly swallowed by a lookup that matched closed pull requests. Both are now handled.
  • A sleeping laptop looks exactly like a hung harness from the inside. During an unattended suite the lid was closed, the cluster froze with the machine, and on every wake the sandbox reaper saw fifteen idle minutes and deleted the running sandboxes. Two tasks failed. The caps cannot see a paused world; the reaper could. Keep the machine awake.

Decisions I would defend

  • One check script. The same script runs on a laptop, in the sandbox and in CI, so "passes checks" means the same thing everywhere. A project in another language brings its own check script and the root one calls it.
  • Gateway aliases. Code calls primary; the real model and keys are configuration. Switching models was an environment change and a redeploy.
  • Standalone sample projects. The sandbox contains only what the agent works on. The harness lives in a separate workspace and never enters the sandbox.
  • Eval runs skip the pull request. They are scored by hidden tests in the sandbox, so the gate measures the agent, not GitHub plumbing.
  • The console is a page over the API. No second service, no build step, no second source of truth. Time series stay in Grafana.

What I would do next

A second model provider behind the fallback alias, so a provider outage is covered instead of merely observed. The GitHub App, so agent pull requests carry the bot's identity rather than a human's. Branch protection and a self-hosted runner for the eval gate. Per-user identity on the console. And a consumer for the dataset export, because the flywheel is built and nothing is drinking from it yet.


Everything in the project uses invented data. Images, charts and dependencies are pinned. Secrets live in an env file and Kubernetes secrets and never in the repository, the logs, the transcripts or the dataset.

Written and built by Shankar Prabhu. Diagrams are inline SVG rendered from Mermaid sources; the page has no scripts beyond the framework and no tracking.