ACTIVATED HUMAN/ ai

How do you see what an AI agent did and why?

Record a trace for every run: each model call, each tool call and its result, what the agent read, what it decided, and what it cost, linked together in order. With traces you can open any run and find the step that went wrong, instead of guessing from the final answer.

Connect the traces to your existing logs and alerts, so agent failures show up where your team already looks.

What a trace should record

A useful trace answers "what happened on this run" without anyone rerunning it. For each run we record:

FieldWhy it matters
The input and who sent itTies a complaint to the exact run behind it.
Context given to the modelShows whether the agent had the facts it needed or was missing them.
Each model callPrompt, response, model version, and latency, so a behavior change can be tied to a model change.
Each tool call and resultShows what the agent did in your systems, including calls that failed.
Decisions and approvalsWhat the agent chose, and whether a person approved it.
Tokens and costCost per run, so you can see which runs are expensive and why.

How agent monitoring differs from service monitoring

Standard service monitoring asks whether the service is up, how fast it responds, and how often it errors. An agent can pass all three and still be wrong: it responds quickly, returns no errors, and tells a customer something false. So agent monitoring has to capture the content of each decision, as well as the timing.

It also has to capture what the agent knew at the time. When an agent with memory makes a bad call, the question is often which stored fact it relied on and when that fact was written, so a trace should record which memories were retrieved on each run. The trace sits next to the ordinary service data, so a slow database or a failing API shows up beside the agent run it affected.

Using traces to fix problems

When someone reports a bad result, we open the run, find the step where it went wrong, and fix the cause: missing context, a tool returning bad data, an instruction the model misread, or a gap in the agent's permissions. The failing run then becomes a new eval case, so the next release is tested against it.

Traces also drive alerts. We set alerts on rising tool failure rates, runs that exceed their cost budget, and runs that loop past a step limit, so a problem pages a person before a customer notices it. In our own operations, production errors and pager incidents become tasks for agents, which come back as reviewed code changes.

What we set up

In our own stack, traces go to Grafana Tempo beside our logs in Loki, errors go to Sentry, and PagerDuty pages a person when something breaks. Every model call is also logged with its tokens and cache use into our analytics database, so cost can be broken down per message.

For a client, we pick hosted or self-hosted tools depending on where your data may be stored, and we set rules for what is kept. Personal data can be masked in traces, and retention can be limited to what your policies allow. You get dashboards your team can read without us, and the tracing code lives in your repository.

Getting this set up

If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.

Working with us
AI systems audit$6,500One week
We spend one week with the people doing the work.
Workflow agentsFrom $12,0002 to 4 weeks
An agent that takes over one recurring job and does it on its own in production, connected to your tools, with evals, tracing, and spending limits.
Agent systemsFrom $30,0004 to 8 weeks
Agent systems that run a core part of your business or your product in production, on your cloud or ours, with a custom harness, evals, tracing, and spending limits.

Questions

What is the difference between AI agent observability and LLM monitoring?

LLM monitoring usually tracks single model calls: latency, tokens, and errors. Agent observability links every model call, tool call, and decision in a run together, so you can see how the agent reached its result and which step caused a problem.

Can traces include sensitive customer data?

They can, and they should be handled like any other copy of that data. We mask personal fields where possible, store traces where your policies allow, and set retention limits.

Do we need to change our logging setup to trace an agent?

Usually not. Traces can go into the tools you already use or into a separate store linked to them. Your team can then check agent runs and service errors in the same place.

Related questions