ACTIVATED HUMAN/ ai

How do you change an AI agent without breaking it?

Handle every prompt, model, and tool change like a code change: keep it in version control, review it, and run it against the full eval suite before it ships. Add each production failure to that suite, so a fix made once is checked on every change after it.

Where ordinary code can catch a failure, add a code check as well, because prompt instructions can stop working when the model changes.

Why agent changes break things quietly

A change to a prompt or model can shift decisions without changing how the conversation reads. After a change, an agent might approve something it used to send to a person, classify a case differently, or skip an escalation. The replies still look reasonable, so nobody notices until a customer or an auditor does.

Changes also arrive from outside. Providers release new model versions and retire old ones, and a prompt tuned for one version can behave differently on the next. Tools change as well: a vendor renames a field, and the agent starts filling it with the wrong value. Each of these needs the same protection, which is a test suite that checks decisions and actions on real cases every time anything changes.

The release process we set up

  1. The change is made in version control: prompt text, model choice, tool definitions, and limits all live in the repository.
  2. The full eval suite runs against the change, including every past failure kept as a regression case.
  3. Decisions on the same cases are compared with the current version, so any case that now gets a different answer is reviewed by a person.
  4. The change runs in a separate preview environment with its own data before it reaches production.
  5. After release, traces and alerts are watched for changes in failure rate, cost, and escalations.
  6. Rollback means pointing the deploy back at the previous version in Git.

In our own development, a merge request can get its own preview environment with its own databases and a public address, removed when the change merges.

If a coding agent wrote your product and every fix now breaks something else, our technical partner service puts this kind of gate in front of every deploy and runs it for you.

When a model version changes

Pin model versions, so an agent does not change behavior because a provider updated a default. When a new version is released, we run the same eval suite against it and compare quality, cost, and speed on your cases. Only then do we decide, per task, whether to switch.

This also answers a common question: when to adopt a newer, more capable model. A newer model can score higher on public benchmarks and lower on your cases, so we switch only when it passes your eval suite, and we write down the comparison so you can see why a model was chosen.

Fixing with code where code can do it

Adding a sentence to the prompt is the quickest fix, but it can stop working when the model changes, and each added rule competes with the others for the model's attention. When a failure can be detected by ordinary code, we check it with code.

For example, our evals showed a model pairing dates with the wrong weekday, so code now checks every date and weekday pair in that product's output before it reaches a user. Permissions work the same way. The agent cannot perform a write action until a person approves it, and that rule is enforced by the tool layer, whatever the prompt says.

Getting this set up

If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.

Working with us
AI systems audit$6,500One week
We spend one week with the people doing the work.
Workflow agentsFrom $12,0002 to 4 weeks
An agent that takes over one recurring job and does it on its own in production, connected to your tools, with evals, tracing, and spending limits.
Agent systemsFrom $30,0004 to 8 weeks
Agent systems that run a core part of your business or your product in production, on your cloud or ours, with a custom harness, evals, tracing, and spending limits.

Questions

How do you stop old agent failures from coming back?

Every fixed failure becomes a permanent test case in the eval suite, and the suite runs before every change ships. If a later prompt or model change brings the failure back, the release fails before it reaches production.

How do you catch a prompt that breaks when the model version changes?

Pin the model version, and run the full eval suite against any new version before switching. Compare decisions case by case with the current version, and review every case where the answer changed.

Should prompts be stored in code or in a prompt management tool?

Either works if prompts are versioned, reviewed, and tested before release. We keep them in the repository by default, so a prompt change goes through the same review and tests as any other change.

Related questions