Running AI agents in production
Running an AI agent in production means it does real work every day, on real data, and you can check that the work was done correctly. A working demo can take days to build. Keeping the agent reliable after launch takes tests on every change, a record of every step, spending limits, and connections to the systems it depends on.
Each section below links to a longer answer. It is based on how we build and run agents for our own products and our own operations.
Why agents that work in a demo fail in production
A demo runs on clean inputs with a person watching. In production the data is messy, tools time out, and people ask for things nobody tested. Small error rates also compound: an agent that gets each step right 95 percent of the time finishes a ten-step task without a mistake only about 60 percent of the time.
Costs change too: retries and loops that go unnoticed in a demo repeat on every production request. Most of these failures are fixed outside the model, in what context it receives, which tools it may call, what it may change without a person approving, and a record of each run.
Knowing the agent is right: evals
An eval is a repeatable test for an agent: a set of real inputs, the result each one should produce, and a check that runs automatically. We build evals from your own work, such as past tickets, documents, and cases where a person had to step in. We add cases designed to make the agent fail, such as instructions hidden inside a document, requests it should refuse, and questions its data cannot answer.
Some checks are exact: did the refund call succeed, is the date right, did the agent call the correct tool. Others need judgment, and a second model grades them against written criteria. Exact checks run beside the graded ones, so a well-written reply cannot hide a failed action. The full suite runs before every change ships.
Seeing what the agent did: traces
A trace is the full record of one agent run: the input, what the agent read, each tool call and its result, each model call, the decision it made, and what the run cost. With traces, you can open the run behind a complaint like "it gave the customer the wrong date" and see which step went wrong.
We set up tracing from the first day of a build and connect it to your existing logs and error reports, so an agent failure appears next to the service failures around it. Traces also feed the evals. A bad run in production becomes a new test case, and the next release is checked against it.
Changing the agent without breaking it
A reworded prompt, a new model version, or a new tool can each change what the agent decides. The risk is a change that fixes one case and quietly breaks another, such as an agent that now approves something it used to send to a person.
We handle prompts, model choices, and tool definitions like code. They live in version control, go through review, and run against the full eval suite before they ship. When a provider releases a new model version, the same suite runs against it before anything switches. When we fix a production failure, the case that caused it joins the suite. Where ordinary code can catch a failure, we add a code check instead of more prompt instructions. For example, our evals caught a model pairing dates with the wrong weekday, so code now checks every date and weekday pair in that product's output.
Context and state
Context is everything the model sees on a given call: instructions, the conversation so far, retrieved documents, and tool results. With too little, the agent guesses. With too much, it gets slower, costs more, and misses the detail that matters. Context management decides what goes in, in what order, and what gets summarized or dropped as a task runs long.
State is what the system knows about a task outside the model: which steps are done, what was approved, and what is waiting on a person. We keep state in a database and in durable workflows, so a task survives a restart, a deploy, or a provider outage and continues where it stopped. We have built the tool-calling and context layer for an agent that answers cost questions across dozens of billing and usage data sources on AWS, Google Cloud, and Azure.
Memory that holds up over months
Memory is what an agent carries from one conversation or task to the next: a customer's history, a decision made last month, the names your team uses for things. The hard problems show up after months of use, when some stored facts are out of date, two of them contradict each other, or one person's information must never reach another.
We build memory as a data system with rules. Each person or customer gets a separate search index, and older history is summarized instead of stored word for word. In the systems we build for you, each stored fact also keeps its source and date so the system can tell which one is current, and permissions apply to memory the same way they apply to documents. Our own products carry months of history per user with separate indexes and layered summaries.
How do you give an AI agent memory that holds up over months?
Connecting to tools with no official integration
Most businesses run on software that was never built for AI: an agency management system, an estimating program, a vendor portal, a shared spreadsheet. If a system has an API, sends webhooks, exports files, or has a web page a person can use, an agent can usually work with it.
We build each connector as ordinary code with a narrow job, and the agent calls it as a tool. Every tool has limits on how often it runs and what it can change. Actions that write data, send messages, or move money can wait for a person to approve them, and the approved action then runs exactly as it was shown. API keys and passwords stay outside the agent's environment. Our operations agents are fed by 35 webhook routes from seven sources, and every request is checked for a signature. One of those services offers no way to sign its requests, so we wrote a signing scheme for it.
Choosing models
No single model is best for every task. A large model can be worth its price for a judgment call and wasteful for sorting an inbox. We choose per task, based on how each model scores on your evals, what it costs per run, and where your data is allowed to go. Claude, OpenAI, Gemini, and open-source models are all candidates, and many systems use more than one.
We run several model providers in production ourselves, with automatic fallback when one has an outage. For sensitive work we set a minimum model level that cost savings can never go below. When we recommend a vendor, the eval results behind the recommendation come with it.
Keeping costs under control
Most of an agent's cost comes from four things: long context resent on every call, retries, loops, and a large model used where a small one would do. We measure spend per run, per task, and per customer, down to the token and to whether the provider's prompt cache was used.
Prompt caching makes repeated context much cheaper to resend, but only when prompts are built so the repeated part stays identical from call to call. We set hard limits on every agent and every tool, so a runaway loop stops at a budget instead of running overnight. When we rebuilt cost reporting for our own products, we found that our monitoring vendor's default math overstated the cost of cache-heavy calls by about six times, so we check reported costs against the provider's own pricing.
Infrastructure
An agent in production needs what any other production service needs: somewhere to run, a database for its state, a workflow engine for long tasks, logs, backups, and a person who gets paged when it breaks. We set this up on servers you already have, a low-cost VPS, a new AWS, Google Cloud, or Azure account in your name, or hardware we own. You hold the keys in every case.
Long-running tasks run on a durable workflow engine, so a task that waits hours for an approval does not fail when a server restarts. We run our own products on three Kubernetes clusters, with deploys through Git, backups we have restored from, and on-call paging. A single agent on one small server gets the same deploys, backups, and paging.
The agent harness
The harness is the code around a model that turns it into an agent: the loop that calls the model, the tools it may use, the context it receives, the limits on what it may do, and the record of each run. Frameworks such as the Claude Agent SDK and LangGraph provide a general harness, and they are the right starting point for many jobs.
A custom harness is worth building when the general one gets in the way, for example when approvals must be enforced at the tool level, each tool needs its own budget, or state must survive restarts. The harness behind our largest product has 127 registered tools, each with its own execution budget, and write actions that run only after approval.
Agent graphs
An agent graph splits a job into steps with defined paths between them, such as research, draft, check, and send, with branches for exceptions and a point where a person approves. Each step can use its own model, instructions, and tools, and each step can be tested on its own.
Graphs make long jobs predictable, because the order of work is set in code and the model makes decisions only inside each step. They also make failures easier to find, since a trace shows exactly which step produced a bad result. Graph failures often happen at the handoffs between steps, where one step's output does not match what the next step expects, so in the graphs we build each handoff is validated against a defined format.
Where to start
If you have an agent that works in testing and struggles in production, or a job you want an agent to take over, start with the one-week audit. In one week we rank the work worth automating by the hours it gives back, and you have a first automation or a working proof of concept by Friday. Larger builds are quoted in writing after the audit.
Questions
Why do AI agents fail in production when they worked in the demo?
A demo runs on clean inputs with a person watching. Production adds messy data, slow or failing tools, and requests nobody tested, and small error rates compound across many steps. The usual fixes are evals built from real cases, a trace of every run, limits on what the agent can change, and spending caps.
How long does it take to put an AI agent into production?
An agent that takes over one recurring job takes 2 to 4 weeks to build. A system that runs a core part of a business takes 4 to 8 weeks. When the job is not yet defined, the one-week audit comes first and ends with something working.
Do we need our own engineers to run an AI agent?
No. We build it with tests, traces, and alerts, and we can run it for you. If you have engineers, we hand over the code and documentation and work alongside them until they are comfortable running it.
Who owns the agent you build?
You do: the code, the accounts, and the documentation. It runs in your cloud, on a server you give us, or on our infrastructure, and it keeps running without us.
- How do you test an AI agent?
- How do you see what an AI agent did and why?
- How do you change an AI agent without breaking it?
- Why do AI agents lose track of long tasks?
- How do you give an AI agent memory that holds up over months?
- Can an AI agent work with software that has no API?
- Which AI model should a business use?
- How do you keep AI agent costs under control?
- What is an agent harness, and do you need a custom one?
- What is an agent graph, and when do you need one?
