ACTIVATED HUMAN/ ai

Is my AI-built app really finished?

Probably not all of it, and the parts that are not finished look exactly like the parts that are. A coding agent builds what satisfies the request in front of it, and the fastest thing that satisfies "add a reports page" is a page that shows a report. Whether the report is calculated from your data, filled with sample numbers, or wired to nothing yet is invisible from the screen. The agent will also call it done, because from where it sits, the request was met.

What you need is an inventory: every feature, marked real, partly built, or screen only, with the evidence. You can get most of the way there in an afternoon without reading code, by using the app the way a stranger would and by asking the agent the right questions in writing. The rest takes someone who reads the whole codebase.

This page lists the ways an app can be unfinished while looking finished, how to check each one on your own app, and what a full read produces.

In founders’ words

“And once you had a few different feature areas going, did you start losing track of your own decisions? Do you find yourself re-explaining the same context every time you open it? Has it ever just quietly built something different from what you actually meant?”

r/startups, June 2026 · source

“The customer did everything right. She filled in all the information, even more than required. Somewhere in my system, something silently broke.”

r/vibecoding, December 2025 · source

“I had Cursor write the endpoint that loads a user's saved views, looked fine in the diff so I approved it and moved on.”

r/cursor, October 2026 · source

The ways an app looks done and is not

One of these is never the whole story. Go through all of them.

  • Screens with nothing behind them. A button, a form or a settings page that renders, and saves nothing, sends nothing, or changes nothing.
  • Half-connected flows. The form saves, but nothing reads the record back; the email is written, but never sent; the upload lands, but no page shows it.
  • Placeholders standing in for logic. A total that is a fixed number, a list that is sample data, a chart drawn from a hard-coded array, a comment that says what the code would do.
  • Fake success. The screen says "Sent", "Saved" or "Paid" whatever happened underneath, because the message was written before the call was.
  • Only the happy path. It works when every field is filled, the network is up and the payment succeeds, and has no answer for anything else.
  • Stand-ins for outside services. The payment, email, maps or AI call returns a sample response because no real key was ever set, and nobody noticed since the sample looked right.
  • Rules that live only in the screen. A paid-only page that is hidden rather than locked; an admin page that any signed-in user can open by typing its address.
  • Tests that test nothing. A test suite that passes because the tests check that the code ran, not that it did the right thing.

Why the agent does not tell you

The agent is not hiding anything. Each request is answered on its own, and "it works in the preview" is a complete answer to most of them. When the request was "add a reports page", a page with plausible numbers is a correct result. Nobody asked for the numbers to be calculated, because to someone who is not an engineer, that part is obviously included. To the agent it is a separate job that nobody named.

The gap is hidden from you for a second reason: you do not know what to ask. An engineer looking at the same app asks where the total comes from, what happens on a failed payment, and which key the email service uses. Those questions come from having run software, not from reading the screen. That is the whole difficulty, and it is not a failing on your part.

Checks you can do without reading code

  • Do everything for real, once. Sign up as a new user in a private window, with a real email you can read. Create a record, close the browser, come back, and look for it. Trigger every email and open your inbox. Pay with a test card and then look for the payment in your payment provider's dashboard, not in your app.
  • Look at the data, not the screen. Open your database dashboard after each action and find the row. If the screen shows a number, find where that number is stored or ask what it is computed from. A number with no row behind it is a placeholder.
  • Break the happy path on purpose. Leave a field blank, submit twice, upload the wrong kind of file, turn off wifi mid-save, use a card that declines. Watch for "Saved" or "Sent" messages that appear anyway.
  • Check every outside service. List every service the app is supposed to talk to (payments, email, AI, maps, storage). For each one, ask the agent whether a real key is set and where, then make the app use the service and look for the result on the provider's side.
  • Ask the agent for an inventory, in writing, with no code changes. Every screen and every button, what each one calls, and what it writes. Then ask it to search the code for sample data, placeholders, TODO, mock, "not implemented" and "would", and to list each hit with the feature it belongs to.
  • Ask for a status per feature. Real, partly built, or screen only, with one line of evidence each. Then test three of the ones it marked real. If one fails, the list is not trustworthy and the read has to be done by a person.

What a full read produces

When someone who has run software reads the whole codebase, the result is a short document, not a list of bugs. It says which features are finished and tested, which are only partly connected and where the gap is, which are only screens, what will break first when real customers arrive, and what to do in which order. It also says what to leave alone, because untangling the wrong thing wastes a week.

That document is what you hand to anyone who helps you next, and it is how you stop paying for the same feature twice. It is the shape of what we produce in a first look. Whether or not you work with us, insist on that shape from whoever reads your code: statuses with evidence, not a quote.

When you do not need help, and when you do

If nobody pays yet and the app is for you or a few friends, the checks above are enough. Work through them, keep the list, and fix the screen-only parts one at a time with a test for each. You will know your own product far better at the end.

Bring someone in when customers pay or are about to, when the app holds other people's data, when the agent's own inventory failed the spot check, or when every fix now breaks something else. At that point guessing is the expensive option. We start with a free call, and on it we will tell you if the checks above are all you need.

What you see, what may be underneath, and how to check it yourself
What you seeWhat may be underneathHow to check
A dashboard with numbersSample data or a fixed figureAdd a real record and see whether the number moves
"Email sent"A message written before the call, or a service with no real keyTrigger it and open the inbox; ask where the email key is set
"Payment successful"The screen decides, not the serverPay with a test card, then find the payment in the provider's dashboard
A settings pageValues that are saved and never readChange a setting and look for the effect somewhere else in the app
An admin areaHidden from the menu, open to anyone who types the addressOpen it as an ordinary user
"All tests pass"Tests that check the code ran, not what it didAsk for the test names and match them to the cases you care about
A feature the agent marked doneThe request was met; the system behind it was not builtDo it for real, then look for the row in the database

Questions

The agent says it is done. Why would it be wrong?

Because "done" means the request in front of it was satisfied, and a page with plausible numbers satisfies "add a reports page". The parts you assumed were included, the calculation, the error handling, the real connection to the email service, were separate jobs nobody named. The agent is reporting what it did, not what you meant.

Can I just ask the agent to list what is real and what is not?

Yes, and you should, with no code changes allowed during the answer. Then test three of the things it marked real. Agents grade their own work generously, and a list you have spot-checked is worth far more than one you have not.

Is this because I did not write good enough prompts?

No. The missing parts are the ones an engineer knows to ask about and a non-engineer reasonably assumes are included. Better prompts help at the edges. What closes the gap is a written spec of the cases before each feature and tests that fail until the real thing works.

Should I rebuild from scratch once I find out how much is not real?

Rarely. The screens, the data model and the parts that are real are worth keeping, and a rebuild with the same unclear spec produces the same gaps faster. Finish the partly built parts one at a time, behind tests, in the order the inventory says matters.

How long does it take to find out?

The checks on this page take an afternoon for most apps. A full read by a person depends on the size of the code and how much of it is only screens; nobody can quote that honestly before reading it. Be careful of anyone who can.

Working with us
  1. First look, $750. After a free call, we read your whole product and tell you what is finished, what is not, and what to do first.

  2. Setup, $3,000 fixed. We make it ready for real customers, in accounts you own.

  3. Partner, $2,500 a month. We review what your coding agent writes and keep the checks and tests current. Month to month.

Related questions