ACTIVATED HUMAN/ ai

My app's AI feature got worse after an update. What happened?

Something changed between the model and your users, and nothing was in place to notice. There are three usual causes: the provider updated or retired the model behind your feature, your coding agent edited the prompt while fixing something else, or the data the feature reads changed shape. Each one changes answers quietly, and the first person to notice is a customer.

The fix for this time is to find which of the three it was and put it back. The fix for every time after is a set of evals: real examples from your product with the answer you expect, rerun after every prompt or model change, so you see whether answers got worse before customers do. We run evals on the AI features in our own products for this reason.

If the AI feature is a nice-to-have and nobody has complained, you may not need help yet. Build the example set below. It takes an afternoon.

In founders’ words

“Anyway every app calling this model by API with a good prompt system, will have their agent changing a lot. In my case I am totally sure I will loose the effect I am looking for.”

OpenAI developer forum, November 2025, November 2025 · source

“Today, while running acceptance testing, it decided to stop conforming to the schema 1 out of every 3 requests. To fix it, I tweaked the prompts.”

r/ArtificialInteligence, October 2025, October 2025 · source

“Anyways, agents drift. You have to keep telling them who they are.”

r/LLMDevs, July 2026, July 2026 · source

The three usual causes

  • The model changed. Providers release new versions under the same name and retire old ones on a schedule. A model name without a date is a moving target. A dated snapshot stays the same until the provider retires it, and they announce that in advance.
  • The prompt changed. Coding agents rewrite prompts when you ask them to fix something nearby. "Make the summary shorter" can quietly remove the line that told the model to answer in the customer's language. The change is in your code history, if the prompt lives in a file.
  • The input changed. A new field, longer documents, empty values, a different date format. The prompt is the same, the model is the same, and the answers are worse because what goes in is different.

Settings count as prompt changes: temperature, maximum length, which tools the model may call. Agents adjust these while fixing other things.

What an eval is, in plain words

An eval is a list of real inputs from your product, each with what a good answer looks like, and a script that runs all of them through your feature and scores the results. Where the answer can be checked exactly, the script checks it: the right category, valid JSON, the customer's name present. Where it cannot, a second model grades the answer against your description of a good one.

You run it after every change to a prompt, a model or a setting, and compare the score to the last run. If it drops, you do not ship the change. Start with twenty to fifty examples taken from real use, and add every case a customer complains about. Keep the examples in the repository next to the prompt.

Your coding agent can write the runner. The examples and the expected answers have to come from you, because you are the one who knows what good looks like.

Finding out which change it was today

  • Collect five of the bad answers with their exact inputs. Ask the customers, or pull them from your logs if you keep them.
  • Look at the code history for the prompt file and the settings around it. If anything changed in the last few weeks, that is the first suspect. Put the old version back and run the five inputs again.
  • Check the provider's model page. If your code names a model without a date, the provider may have moved it. If they sent a retirement notice, the replacement is a different model with different habits.
  • Compare this week's inputs with last month's. Longer, emptier or differently shaped inputs point at the data.
  • Whichever one you find, make the five inputs the first entries in your eval set. That is how the set starts.

Keeping it from happening again

Pin the model by its dated name and put the retirement date on your calendar. Keep the prompt in its own file with a version number, so a change is visible in the history and you can go back. Log every request and answer, with private data removed, so you can replay a complaint. Configure a second model as a fallback, so an outage at one provider degrades the feature instead of removing it.

Then make the evals run automatically whenever code changes, and have them block the change if the score falls. Once that is in place, the agent can edit prompts freely, because the evals catch what it breaks.

When the model itself got worse

Sometimes you changed nothing and the answers still changed. The provider updated the model, or retired the one you used. You have three options: go back to a dated snapshot if one still exists, move to another model after running your evals on it, or adjust the prompt until the evals pass again. The evals are the judge in all three. Without them you are choosing on "it looks better", and that is how the feature got worse in the first place.

If your product depends on the feature and the provider has announced a retirement, bring someone in before the date, not after. Moving a feature to a new model with evals in hand is a day of work. Doing it live, from complaints, is a bad week.

Which change was it?
What changedHow to tellWhat to do
The prompt or its settingsCode history shows an edit to the prompt file, temperature or lengthPut the old version back. Add the bad inputs to the eval set.
The model, under the same nameYour code names a model with no date; the provider's changelog shows an updatePin a dated snapshot. Run the evals on the new version before moving to it.
The model was retiredA notice from the provider, or errors in the logsRun the evals on two or three candidates. Pick on score, not on the newest name.
The input dataThis week's inputs are longer, emptier or shaped differentlyHandle the new shape in code. Add examples of it to the eval set.
Nothing you can findSame prompt, same dated model, same inputs, different answersLower the temperature, add exact checks to the evals, and measure how often it varies.

Questions

Which model should I pin?

The dated version of the one your evals pass on. Write its retirement date down when the provider publishes it. Pinning the newest model without running evals on it is the same gamble as not pinning.

Can my coding agent write the evals?

It can write the runner and the exact checks. The examples and what a good answer looks like need to come from you or from real customer cases. An agent writing its own examples tests what it already assumes.

How many examples do I need?

Twenty to fifty to start, taken from real use and including the ones that went wrong. The set grows by one each time a customer reports a bad answer.

Does running evals cost money?

Each run calls the model once per example, plus a grading call for the fuzzy ones. For fifty examples that is small change, and far less than one customer leaving over a bad answer.

Can I just switch to the newest model and move on?

Only after the evals pass on it. Newer models are better on average and different in particulars, and your prompt was written for the old particulars.

Working with us
  1. First look, $750. After a free call, we read your whole product and tell you what is finished, what is not, and what to do first.

  2. Setup, $3,000 fixed. We make it ready for real customers, in accounts you own.

  3. Partner, $2,500 a month. We review what your coding agent writes and keep the checks and tests current. Month to month.

Related questions