What is an agent graph, and when do you need one?
An agent graph breaks a job into steps with defined paths between them, such as research, draft, check, and send, with branches for exceptions and stops for a person's approval. You need one when a job has more than a few steps, when different steps need different models or tools, or when you must be able to show exactly how a result was produced.
Single agent or graph
| Question | Single agent | Agent graph |
|---|---|---|
| Who decides the order of work? | The model, as it goes | Code, with the model deciding inside each step |
| How predictable is it? | Varies from run to run | Same path for the same kind of case |
| Can steps use different models? | One model throughout | Each step uses the model it needs |
| Can each step be tested alone? | Hard | Yes |
| Best for | Open-ended questions and short tasks | Repeatable jobs with several stages |
Designing the graph
We start from how the job is done today. Each stage a person goes through becomes a step, and each point where they check something or ask someone becomes a check or an approval. The parts that always happen in the same order are set in code, and the model decides only what needs judgment inside a step.
Checks sit between steps. A draft can be reviewed by other agents looking at it from different angles, such as accuracy against the source, tone, and policy, before anything goes out. A step that sends a message, changes a record, or moves money waits for a person when the risk calls for it. In our own operations, scheduled and event-driven agents do the daily work, and a person decides anything that matters.
Where graphs fail in production
A common graph failure happens at the boundary between steps. One step returns a field in a slightly different format, leaves a value empty, or says "none found" in prose instead of an empty list, and the next step acts on it. In the graphs we build, each handoff is validated against a defined format, and a mismatch is treated as a failure with a limited number of retries.
Loops through outside systems are the other common failure. An agent updates a ticket, the ticket system sends a webhook, and the webhook wakes the agent again. We built a breaker that stops such loops, and a filter that collapses bursts of events: in our own operations a burst of 40 events becomes one task carrying the latest state.
Running graphs durably
A graph that runs for minutes can live in memory. One that waits hours for an approval, or runs every night across thousands of records, needs a durable workflow engine that records each finished step and resumes after a restart. Our largest product runs 23 durable workflows on Temporal, and each one picks up where it stopped after a deploy or outage.
Every step of every run is traced, so when a graph produces a bad result, the trace shows which step did it and what that step was given.
Getting this set up
If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.
Questions
Is LangGraph ready for production use?
It can be, when you add what production needs around it: durable state, evals, tracing, limits, and validation at every handoff between steps.
What is multi-agent orchestration?
Coordinating several agents on one job, each with its own role, tools, and model, with code deciding how work passes between them. An agent graph is one way to organize it.
Do agent graphs need a human in the loop?
For steps that send messages, change records, or move money, yes, at least until the step has a production record. Approval points are part of the graph design, and they can be loosened for a step once its eval results and production record justify it.
