ACTIVATED HUMAN/ ai

How do you keep AI agent costs under control?

Measure cost per run and per task, cache repeated context, route simple steps to cheaper models, and put hard spending limits on every agent and tool. Most runaway bills have a few causes: long context resent on every call, retries and loops, and a large model used where a small one would do.

Where agent costs come from

  • Context resent on every call: instructions, history, and documents billed again each time.
  • Long conversations and tasks, where every step carries the whole history so far.
  • Retries and loops: a tool keeps failing and the agent keeps trying.
  • The most expensive model used for every step, including simple ones.
  • Tool and data costs outside the model: search, databases, browser sessions, and paid APIs.
  • Agents calling other agents, where one request turns into dozens of model calls.

Measuring cost per run, task, and customer

A monthly provider invoice cannot tell you which agent, task, or customer drove the bill. We record every model call with its tokens, its cache use, and the task and customer it served, so cost can be broken down any of those ways. That shows which workflow is expensive and whether a change made it cheaper.

Reported costs can be wrong. When we rebuilt cost reporting for our own products, we found that our monitoring vendor's default pricing overstated the cost of cache-heavy calls by about six times. We now calculate cost from the provider's published prices ourselves instead of trusting the vendor's defaults.

If you built a product with a coding agent and want someone to run it and tell you each week what it cost, see our technical partner service.

Prompt caching

Providers charge much less for input they have recently seen, but only when the start of the prompt is identical, byte for byte, to a recent call. A timestamp or user name near the top of a prompt breaks the cache on every call.

So we split each prompt into a fixed part that never changes (instructions, tool definitions, reference material) and a changing part that goes at the end. In our own products, a large batch of work starts at 4 AM each day, and just before it we send one request per distinct fixed part so the cache is ready when the batch arrives. Caching lowers cost without changing what the model produces.

Hard limits that stop spending

Alerts tell you about overspending after it happens, so we also set limits that stop a run when it reaches its budget. Every tool in our own harness has its own execution budget, agents can hand work to other agents only two levels deep, and each product tier has its own spending limits for expensive features like voice.

For a client, we set limits per run, per task, and per day, and decide in advance what happens when one is reached: stop, fall back to a cheaper model, or ask a person. With these limits in place and routine steps on smaller models, monthly cost has a known ceiling.

Getting this set up

If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.

Working with us
AI systems audit$6,500One week
We spend one week with the people doing the work.
Workflow agentsFrom $12,0002 to 4 weeks
An agent that takes over one recurring job and does it on its own in production, connected to your tools, with evals, tracing, and spending limits.
Agent systemsFrom $30,0004 to 8 weeks
Agent systems that run a core part of your business or your product in production, on your cloud or ours, with a custom harness, evals, tracing, and spending limits.

Questions

Why is our AI API bill so high?

Usually because long context is resent on every call without caching, retries or loops run unchecked, or a large model handles steps a small one could. Measuring cost per run and per task shows which of these is driving your bill.

How do you attribute AI costs to agents, teams, or customers?

Tag every model call with the agent, task, and customer it served, and record its tokens and cache use. Costs can then be totaled any of those ways, including across agents that call other agents.

Does cutting AI costs reduce quality?

Caching and better context selection lower cost with no change in quality. Moving a step to a cheaper model can affect quality, so we make that change only when the cheaper model passes the same eval suite.

Related questions