How do you see what an AI agent did and why?
Record a trace for every run: each model call, each tool call and its result, what the agent read, what it decided, and what it cost, linked together in order. With traces you can open any run and find the step that went wrong, instead of guessing from the final answer.
Connect the traces to your existing logs and alerts, so agent failures show up where your team already looks.
What a trace should record
A useful trace answers "what happened on this run" without anyone rerunning it. For each run we record:
| Field | Why it matters |
|---|---|
| The input and who sent it | Ties a complaint to the exact run behind it. |
| Context given to the model | Shows whether the agent had the facts it needed or was missing them. |
| Each model call | Prompt, response, model version, and latency, so a behavior change can be tied to a model change. |
| Each tool call and result | Shows what the agent did in your systems, including calls that failed. |
| Decisions and approvals | What the agent chose, and whether a person approved it. |
| Tokens and cost | Cost per run, so you can see which runs are expensive and why. |
How agent monitoring differs from service monitoring
Standard service monitoring asks whether the service is up, how fast it responds, and how often it errors. An agent can pass all three and still be wrong: it responds quickly, returns no errors, and tells a customer something false. So agent monitoring has to capture the content of each decision, as well as the timing.
It also has to capture what the agent knew at the time. When an agent with memory makes a bad call, the question is often which stored fact it relied on and when that fact was written, so a trace should record which memories were retrieved on each run. The trace sits next to the ordinary service data, so a slow database or a failing API shows up beside the agent run it affected.
Using traces to fix problems
When someone reports a bad result, we open the run, find the step where it went wrong, and fix the cause: missing context, a tool returning bad data, an instruction the model misread, or a gap in the agent's permissions. The failing run then becomes a new eval case, so the next release is tested against it.
Traces also drive alerts. We set alerts on rising tool failure rates, runs that exceed their cost budget, and runs that loop past a step limit, so a problem pages a person before a customer notices it. In our own operations, production errors and pager incidents become tasks for agents, which come back as reviewed code changes.
What we set up
In our own stack, traces go to Grafana Tempo beside our logs in Loki, errors go to Sentry, and PagerDuty pages a person when something breaks. Every model call is also logged with its tokens and cache use into our analytics database, so cost can be broken down per message.
For a client, we pick hosted or self-hosted tools depending on where your data may be stored, and we set rules for what is kept. Personal data can be masked in traces, and retention can be limited to what your policies allow. You get dashboards your team can read without us, and the tracing code lives in your repository.
Getting this set up
If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.
Questions
What is the difference between AI agent observability and LLM monitoring?
LLM monitoring usually tracks single model calls: latency, tokens, and errors. Agent observability links every model call, tool call, and decision in a run together, so you can see how the agent reached its result and which step caused a problem.
Can traces include sensitive customer data?
They can, and they should be handled like any other copy of that data. We mask personal fields where possible, store traces where your policies allow, and set retention limits.
Do we need to change our logging setup to trace an agent?
Usually not. Traces can go into the tools you already use or into a separate store linked to them. Your team can then check agent runs and service errors in the same place.
