Why do AI agents lose track of long tasks?
AI agents lose track when the context they receive is too thin, too long, or out of date, and when the progress of a task exists only inside the conversation. The fix has two parts: context management, which decides what the model sees on each call, and state kept outside the model in a database or workflow engine.
What goes into an agent's context
On every call, the model sees only what it is given. For a working agent that usually includes:
- Standing instructions: the job, the rules, and what the agent must not do.
- The task so far: the request, earlier steps, and their results.
- Retrieved information: documents, records, or past history relevant to this step.
- Tool results: data returned by the systems the agent called.
- Current facts the model cannot know on its own, such as today's date, the user's time zone, and the state of an order.
Each item costs tokens and competes for the model's attention. Context management decides which items go in on each call and in what order.
Why more context is not better
Large context windows make it tempting to send everything. The costs are real: every token is paid for on every call, long prompts are slower, and models get worse at finding the one detail that matters when it sits among thousands of lines that do not.
We give each call a context budget. The agent gets what that step needs, retrieved when needed, and older parts of a long task are summarized in layers, so recent detail stays sharp and older detail stays available in short form. The instructions that never change go first and stay identical from call to call, which also lets the provider's prompt cache cut the cost of resending them.
Keeping state outside the model
The model keeps nothing between calls. If the only record of a task's progress is the conversation, a restart, a timeout, or a long wait for approval can lose the task or run a step twice.
We keep task state in a database and run long tasks on a durable workflow engine, Temporal in our own stack. Each step records that it finished, so after a restart, deploy, or provider outage the task resumes from its last finished step, and steps that send an email or charge a card are written so a retry cannot do it twice. A task waiting days for a person's approval simply waits. Our largest product runs 23 durable workflows this way.
Working across many data sources
Some agents need to answer questions that span dozens of systems. Loading all of that data into context at once does not work. It is too large, and most of it is irrelevant to any given question.
We have built this for an agent that works across dozens of financial and billing data sources, where every number has to be exact and traceable to its source. It calls tools in steps, asking each source a narrow question, and keeps only the results the next step needs, so it can explain a change without anyone digging through consoles and spreadsheets. A business whose answers are spread across a CRM, an accounting system, and a shared drive can use the same approach.
Getting this set up
If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.
Questions
What is context engineering?
Context engineering is the work of deciding what an AI model sees on each call: which instructions, history, documents, and tool results, in what order, within what budget. It usually matters more than how the prompt is worded, because the model can only use what it is given.
Is a bigger context window a fix for an agent that loses track?
Rarely on its own. A bigger window lets you send more, but cost and latency rise and the model can miss the detail that matters. Selecting and summarizing context, and keeping task state outside the model, fixes the cause.
What happens to a long-running agent task when a server restarts?
On a durable workflow engine, the task resumes from its last finished step, as long as each step is written to be safe to retry. Without one, it can be lost or repeat steps, which is a real problem when a step sends a message or moves money.
