ACTIVATED HUMAN/ ai

What is an agent harness, and do you need a custom one?

An agent harness is the code around a model that makes it an agent: the loop that calls the model, the tools it can use, the context it gets, the limits on what it may do, and the record of what it did. Off-the-shelf frameworks are the right start for many jobs. A custom harness is worth building when you need control a framework does not give you.

What a harness does

  • Runs the loop: call the model, run the tools it asks for, return the results, repeat until done.
  • Builds the context for each call within a budget.
  • Defines the tools, their inputs, and what each one is allowed to change.
  • Enforces limits: steps per task, spend per task, and which actions need a person's approval.
  • Keeps state so a task survives restarts.
  • Records a trace of every step.

Framework or custom

You needFrameworkCustom harness
A working agent quicklyGood fitSlower to start
Standard tools and a single modelGood fitNot needed
Approvals enforced at the tool levelOften possible with extra workBuilt in
Budgets per tool and per taskPartialBuilt in
Tasks that run for hours or daysDepends on the frameworkBuilt on a durable workflow engine
Full control of context and cachingLimitedFull

We work with the Claude Agent SDK, LangGraph, and MCP. A build can start on one of them and replace parts as the requirements become clear.

What our own harness enforces

The harness behind our largest product has 127 registered tools, each with its own execution budget. Write tools never run on the model's say-so. When the model calls one, the harness records a proposed action instead. After a person approves it, the recorded call runs exactly as proposed, with no model involved in that step.

Agents can delegate to sub-agents only two levels deep, and sub-agents get read-only access. Context is built for each call within a budget, the fixed part of the prompt is kept identical so the provider's cache applies, and several model providers sit behind it with automatic fallback.

What you get from a custom harness build

You own the harness: the code, the tests, and the documentation, in your repository. It comes with an eval suite built from your work, tracing connected to your monitoring, and limits set for your risk and budget.

For a single recurring job, a framework with a few custom tools and limits is often enough. We add custom parts only where the framework gets in the way.

If your product was built with Claude Code or Cursor and the agent keeps breaking what it just fixed, our technical partner service writes the instructions and review rules it works from.

Getting this set up

If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.

Working with us
AI systems audit$6,500One week
We spend one week with the people doing the work.
Workflow agentsFrom $12,0002 to 4 weeks
An agent that takes over one recurring job and does it on its own in production, connected to your tools, with evals, tracing, and spending limits.
Agent systemsFrom $30,0004 to 8 weeks
Agent systems that run a core part of your business or your product in production, on your cloud or ours, with a custom harness, evals, tracing, and spending limits.

Questions

What is the difference between an agent harness and an agent framework?

A framework is a general library for building agents. A harness is the specific code that runs your agent: its loop, tools, context, limits, and records. A harness can be built on a framework or written from scratch.

Which agent framework should we use?

It depends on the job. We work with the Claude Agent SDK, LangGraph, and MCP, and choose based on your language, hosting, and how much control the job needs over tools, state, and approvals.

How do you stop an agent from taking actions it should not?

Give it narrow tools, limit what each tool can change, and require a person's approval before any write, send, or payment. Enforce those rules in the harness code, so they apply whatever the prompt or the model says.

Related questions