The premise
Most demos of coding agents answer the question "can the model write the code?" That question is mostly settled. The question I wanted to answer is the one an engineering manager asks before letting an agent near a repository: what stops it from doing the wrong thing, and how would I know?
So the harness was designed around three claims, and every component exists to make one of them true in a way you can test:
- Permissions live in code, never in prompts. A prompt is a suggestion. A tool that refuses is a guarantee.
- The agent is checked, never trusted. The agent's own view that it is finished is an input to verification, not a result.
- Every run is a measurement. Cost, outcome, failure mode and transcript are recorded the same way for every run, so failures turn into tests and tests turn into fixes.
A prompt is a suggestion. A tool that refuses is a guarantee.
What it is
A local Kubernetes cluster (k3d) runs the whole thing: a model gateway, the harness API and its console, a warm pool of locked-down sandbox pods, Postgres, ClickHouse and an S3-compatible object store, plus Prometheus, Tempo, Grafana and Langfuse for observability. Argo CD owns the cluster after bootstrap, so every change to it is a commit on main. The agent works on two sample projects in a monorepo: a small FastAPI claims service with a human-owned payments/ module, and a hello-world Next.js app added later to prove the harness is not Python-specific.
The system at a glance
Everything runs in one local cluster, split by namespace. Arrows are the only paths that exist; the sandbox has none out.
A run is a LangGraph state machine with six nodes, checkpointed in Postgres so a capped run can be resumed later in the same sandbox:
plan asks the model for a short numbered plan. implement is the agentic loop, with exactly four tools: list files, read a file, write a file, run the checks. self_check is the agent's side of verification: run every project's lint, types and tests, then review the full diff against the spec. A failure sends the run back to implement with feedback, up to a cap. verify is the harness's own check, run directly in the sandbox, which the agent cannot influence. open_pr takes the staged diff out of the sandbox, re-checks it for regulated paths, and opens a pull request as a bot identity. finalize computes the outcome, writes the transcript, and escalates anything that is not clean.
The console
One page over the API, refreshed every ten seconds. Everything the harness records is here: runs and their outcomes, the specs and which of them may run, the eval suite, the pull requests the loop opened, escalations, the warm pool, and the non-secret configuration.
The stack
Free and open-source editions throughout, every image, chart and dependency pinned, and nothing bespoke where a standard piece would do. The harness itself is a few thousand lines of Python; most of the system is configuration that makes those lines hard to misuse.
Cluster and delivery
- k3d, a local Kubernetes with its own image registry
- Argo CD, app-of-apps from one
deploy/folder; GitOps only after bootstrap - Helm charts, vendor charts for the platform pieces, small in-repo charts for the rest
- GitHub Actions for CI and the eval gate; CODEOWNERS and branch protection as gates
Models
- LiteLLM gateway with two aliases,
primaryandfallback, a global spend cap, and cost in response headers - Claude Sonnet 5 as primary and Haiku 4.5 as fallback, named only in the environment
- Opt-in: the Claude Code CLI in headless mode on a subscription, bridged to the same tools over MCP
Harness
- Python 3.12 in a uv workspace: harness, evals, loop and the sandbox manager as separate packages
- FastAPI and uvicorn for the API; the console is one HTML page served by it, behind an API key
- LangGraph for the run as a state machine, checkpointed in Postgres so capped runs resume
- MCP servers for the tools, in-process; a tool layer with an allow-list and path rules in front of them
- The Kubernetes Python client, used for exactly one verb pair:
pods/execin the sandbox namespace
Sandbox
- One image with both projects, the uv environment, Node 22 and
node_modules, so checks run offline - A warm pool kept by a small manager: claim, touch, reap on idle or age
- Deny-all NetworkPolicy, no service-account token, non-root user, all capabilities dropped, CPU and memory limits
Data
- Postgres for runs, escalations, eval results, failure tags and graph checkpoints
- ClickHouse for the event stream: every tool call, model call, halt and verdict with cost and tokens
- RustFS, an S3-compatible store, for transcripts and the scrubbed dataset exports
- Redis, only because Langfuse needs it
Observability
- OpenTelemetry collector, Tempo for traces, Prometheus for metrics, Grafana for the one dashboard
- Langfuse for a per-call view of what the gateway sent and received
- Every record, event and span carries the run id, the harness version and the model
Quality
- ruff, mypy in strict mode and pytest for Python; ESLint, TypeScript and vitest for the Next.js app
- One
check.shthat discovers each project and runs its checks, identical on a laptop, in the sandbox and in CI - A lock check on every run, so a drifted dependency fails the build
Sample projects
- A FastAPI claims service with SQLite and a human-owned
payments/module - A hello-world Next.js 16 app in TypeScript, added to prove the harness is language-agnostic
Three gates that software cannot pass on its own
The harness refuses to start a run unless the spec is on main with status: approved, and the specs/ folder is code-owned, so approval is a human merge. That is gate one. A pull request needs CI, the eval gate and a human review to merge, and the bot cannot approve its own work. That is gate two. The release environment is a manual Argo CD sync, while the dev environment syncs on every merge. That is gate three. None of these are policies someone has to remember; they are what the system does.
Guarantees and how each one is enforced
The agent cannot touch the regulated payments code
The write tool refuses the path in code and the pull-request tool refuses a diff that touches it. The path list is configuration. The prompt never mentions it.
The agent cannot reach the network or the cluster
Sandbox pods run under a deny-all NetworkPolicy with no service-account token, as a non-root user, with CPU and memory limits. Both projects' toolchains are baked into the image so the checks run offline.
The agent cannot call tools it was not given
An allow-list in the tool layer. A call to anything else is logged as refused.
A run cannot spend without limit
A per-run dollar budget from the gateway's cost headers, a fix-attempt cap, a tool-call cap, a stall watchdog, and a global spend cap at the gateway.
“Done” is not taken at face value
An independent verify step. If the agent stopped and verify fails, the outcome is dishonest_success, escalated, and never exported to the dataset.
Every change to the cluster is attributable
GitOps only, after bootstrap. Argo CD shows the commit and author behind every application.
What a tool call actually passes through
Outcomes are a vocabulary
Every run ends in exactly one of a small set of outcomes, and the whole system speaks in them: the console, the dashboard, the escalation table, the failure tagger.
| Outcome | Meaning | Escalated | Resumable |
|---|---|---|---|
| clean | verified on the first pass | no | no |
| fixed | verified after one or more self-check fixes | no | no |
| dishonest_success | agent stopped, independent verify failed | yes | no |
| capped_budget | cumulative cost passed the budget | yes | yes |
| capped_retries | fix attempts or tool calls hit the cap | yes | yes |
| stalled | no event for the stall window | yes | yes |
| blocked_permission | tried to write under a regulated path | yes | no |
| error | anything else, such as the gateway being down | yes | no |
The loop: a seeded bug finds its own fix
The part I am most pleased with is the improvement loop, because it closed end to end on a real defect. Version 0.1 of the harness shipped with a deliberate bug: the self-check reviewer only read the first 2,000 characters of the diff. Small specs passed. A large spec produced a complete, correct change and then failed self-check three times, because the reviewer saw a truncated diff and called it incomplete, until the run hit its retry cap.
The loop then did its job without a human writing any of it:
- Tagging. Rules map each failed run to a failure mode from its outcome, last error and self-check evidence. The three runs were tagged
truncated_diff. A model is only asked when the rules do not match. - Regression tasks. A failed run becomes a new eval task, with hidden tests, opened as a pull request. Two of the thirteen eval tasks exist because of this.
- A proposed fix. When a failure mode repeats enough times, a proposer drafts a harness change. It may only touch two files, it has to pass the full check script in a scratch clone, it bumps the harness version and the image tag, and it opens a pull request. A human reviewed and merged it. Harness 0.2.0 rolled out through Argo CD, and the large spec passed on re-run.
- A dataset. Verified runs are exported as scrubbed JSONL, with dishonest successes excluded, and a scan proves there are no secrets in it. That export is the handoff point to fine-tuning, which the project deliberately stops short of.
Evals as the gate, not a report
An eval task is a spec plus hidden tests the agent never sees. After a run, the hidden tests are copied into the sandbox and executed, so the score is about the code, not the agent's claims. Each task carries its own test directory and command, which is how a vitest task for the Next.js app scores alongside pytest tasks for the Python service. A pull request that touches the harness, the evals or the loop must meet a pass-rate threshold and must not fall below the stored baseline from main. Other pull requests get a rubric score from a judge model. Scores are stored with the harness version, so before and after a change are always comparable.
Numbers
The dashboard shows cold start with and without the warm pool, agent pull-request merge rate, review minutes per pull request, clean-run rate, cost per run, dishonest successes, failure modes by count and outcomes over time. Those were the metrics the spec asked for, and each has a panel backed by Postgres, ClickHouse or Prometheus. The single most useful number turned out to be the dishonest-success count, because it is the gap between what the agent says and what is true.
Running on a subscription instead of API keys
Late in the project I added an opt-in model backend that runs the agent through the Claude Code CLI in headless mode, billed to a Claude Max subscription, using the CLI's documented long-lived token. The interesting part was keeping the guarantees. The CLI's own tools are disabled, and the harness exposes its four tools to the CLI as an MCP server on loopback. Every call still goes through the same tool layer, so the allow-list, the regulated-path refusal, the event log and the call cap are identical on both backends. The budget works on the cost the CLI reports; a stall, a refused regulated write or the call cap terminates the CLI from outside; and an exhausted usage window becomes a new resumable outcome, capped_rate_limit.
The full suite passed on it. What you lose is real cost accounting, gateway traces and a fallback model, which is why the default stays the gateway.
What went wrong
- A fixed decision had to change. The object store I had planned on withdrew its open-source images. The rule was to stop and ask rather than swap a tool silently, and that is what happened; RustFS took its place.
- Kubernetes exec needs
getas well ascreateonpods/exec, because the websocket upgrade is a GET. The least-privilege role was one verb too strict. - A root-owned mount broke two things at once:
cp -aand git's ownership check inside the sandbox. - The warm pool replaced failed pods forever until it learned to delete them.
- The provider ran out of credits in the middle of a suite. That run shows as 5 of 10 on the dashboard, and it is exactly the outage a second provider is for.
- "Done" as a marker word cost fix attempts. Models sometimes summarise and stop without the magic word. Now, no further tool calls means the agent considers itself finished, and the checks decide whether it is.
- The first fix proposal conflicted after unrelated version bumps, and the second was nearly swallowed by a lookup that matched closed pull requests. Both are now handled.
- A sleeping laptop looks exactly like a hung harness from the inside. During an unattended suite the lid was closed, the cluster froze with the machine, and on every wake the sandbox reaper saw fifteen idle minutes and deleted the running sandboxes. Two tasks failed. The caps cannot see a paused world; the reaper could. Keep the machine awake.
Decisions I would defend
- One check script. The same script runs on a laptop, in the sandbox and in CI, so "passes checks" means the same thing everywhere. A project in another language brings its own check script and the root one calls it.
- Gateway aliases. Code calls
primary; the real model and keys are configuration. Switching models was an environment change and a redeploy. - Standalone sample projects. The sandbox contains only what the agent works on. The harness lives in a separate workspace and never enters the sandbox.
- Eval runs skip the pull request. They are scored by hidden tests in the sandbox, so the gate measures the agent, not GitHub plumbing.
- The console is a page over the API. No second service, no build step, no second source of truth. Time series stay in Grafana.
What I would do next
A second model provider behind the fallback alias, so a provider outage is covered instead of merely observed. The GitHub App, so agent pull requests carry the bot's identity rather than a human's. Branch protection and a self-hosted runner for the eval gate. Per-user identity on the console. And a consumer for the dataset export, because the flywheel is built and nothing is drinking from it yet.
Everything in the project uses invented data. Images, charts and dependencies are pinned. Secrets live in an env file and Kubernetes secrets and never in the repository, the logs, the transcripts or the dataset.