I'm stuck in a test-and-fix loop with my AI-built app
You are stuck in a test-and-fix loop because the agent is working out what the feature should do while it builds it, so every round moves the target. You describe a feature. The agent builds its best guess. You test it, find three things wrong, and describe them. The agent fixes those and shifts something else, because nothing wrote down what "right" was before it started. Add a second agent to test and a third to plan and you have three guesses instead of one, and a bigger bill.
The loop ends when the definition of done lives outside the conversation: a short written spec of the cases, and automated tests that fail until each case works and keep running afterwards. Then the agent's job is to make the tests pass, and a round that breaks an earlier case fails before you ever see it.
This page covers how to write that spec without being an engineer, how to get the tests without writing code, and the signs that the loop is telling you something else.
In founders’ words
“It took 5 full days of back and forth and version numbers hit 0.8.2 and it was just a slog.”
“I currently have a 3 month old inventory project built around my industries workflow and we are in constant testing and amendment loops. I'd like to finish it some time this decade”
Why the loop happens, and why more agents make it worse
A model fills every gap in a request with a plausible default. "Approve orders before dispatch" leaves open who approves, what happens to a rejected order, and whether an order can be edited while it waits. The agent picks answers, builds them, and moves on. When you test and say "no, rejected orders should go back to the store", the agent changes that and, with the earlier choices no longer in front of it, quietly changes a neighbour too.
A tester agent does not help, because it tests what the builder said it built, with the same gaps. A third agent to plan adds a third set of defaults. Each round costs tokens, and five days of rounds is a real bill. The problem is not effort or model quality. It is that "done" was never written down.
Write the spec before the next prompt
Half a page, in your own words, for one feature. The agent reads it before it writes a line. It needs four things:
- The objects. Order, store, approver, status. One line each on what it is.
- Who can do what. "A store user can submit and cancel. A head office user can approve or reject."
- The exact cases, each with a number or a value in it. "An order over the store's limit of 500 waits. An order of 500 or under dispatches at once."
- What a user sees at the end of each case. "The store sees the order marked Waiting, with the reason."
Then ask the agent to criticise the spec and list every question it would otherwise have guessed at. Answer those, and only then let it build. If you run several agents, this is the architect's job: critique, not code.
Tests the agent cannot talk its way past
An automated test is a small program that runs one case from your spec and reports pass or fail. You never read the code; you read the test names, which should match your cases word for word. Ask the agent to write one test per case before it builds the feature, and to show you each test failing first. A test that passes before the feature exists is testing nothing.
Then the rule for the agent is simple: a change is done when every test passes, including every earlier test. That is what stops this round from undoing the last one. For the flows a user clicks through, ask for end-to-end tests, which drive a browser the way a person would. Run all of them on every commit in GitHub Actions or a similar service, so a break fails the build instead of waiting for you.
Habits for each round
- One case per session. The agent makes one test pass, you check it, you save the state.
- Paste the exact error text and the exact steps, not a description. "Clicked Approve on order 1042, saw a blank page, console said: ..." gives the agent something to find.
- After two failed attempts on the same case, go back to the last saved state and make the case smaller. Do not send a third prompt.
- Ask the agent to list the files it changed. If a file outside the feature appears, ask why before you keep it.
- Stop each day with every test green. A loop that runs overnight on a red state is a loop you cannot step out of.
Fewer rounds is also what brings the token bill down. Tests do the tester agent's work for free, every time.
When the loop is telling you something else
If the loop stays on the same area for weeks even with a spec and tests, the feature is probably fighting the structure under it: the same rule lives in several places, or the piece you are building sits on top of something else that keeps changing. The page on every fix breaking something else covers that.
Sometimes the spec itself is the problem: two rules contradict each other, and the agent swings between them. Reading your cases aloud usually finds it.
If the product is a side project and nobody is waiting, the habits above are enough and you do not need help. If a business runs on it and one feature has cost five days and a real bill, a short outside read of the spec, the tests and the code often ends the loop faster than another round. We offer that, and it starts with a free call.
| What you typed | What the agent can build from |
|---|---|
| Head office approves store orders before dispatch | An order over the store's limit goes to Waiting. A head office user sees all Waiting orders and can approve or reject each one with a reason. Approved orders dispatch as before. Rejected orders return to the store as Draft, with the reason shown. |
| Users can export their data | A logged-in user can download a CSV of their own orders only, for a date range they choose. The file has these columns, in this order. If there are no orders in the range, the user sees a message, not an empty file. |
| Fix the stock count, it's wrong | Stock for an item is the opening count plus deliveries minus dispatched orders, counted at the moment the page loads. Orders in Waiting do not reduce stock. Cancelled orders add their quantity back. |
Questions
Do I need a separate testing agent?
No. Automated tests do the testing agent's job, and they do it the same way every time. A tester agent checks the builder's claims with the builder's assumptions, which is how the loop started.
What kind of tests should I ask for if I am not an engineer?
Ask for one test per case in your spec, named in your words, plus end-to-end tests for the flows a user clicks through. Ask to see each test fail before the feature exists. You read the names and the pass/fail list; you never need to read the test code.
The agent says the tests pass but the feature is still broken. Why?
Usually the test checks something other than your case, or it was written to pass. Ask the agent to show the test and explain, in plain words, what it does. If it does not match the case in your spec, have it rewritten and fail first.
How many rounds is normal?
There is no fixed number, but one case that takes more than two attempts is a signal. Either the spec is unclear, or the code underneath is fighting the change. Stop and look at those before a third attempt.
Is there a stack that limits regressions and drift?
The stack matters less than having a repository, a spec, and tests that run on every commit. Within that, pick common, well-documented pieces the model has seen a lot of, because it makes fewer guesses with them.
First look, $750. After a free call, we read your whole product and tell you what is finished, what is not, and what to do first.
Setup, $3,000 fixed. We make it ready for real customers, in accounts you own.
Partner, $2,500 a month. We review what your coding agent writes and keep the checks and tests current. Month to month.
