Implementation worksheet · 6 min read
How to Choose Between Improving and Rebuilding an AI-Generated App
Decide on evidence about where the problems live, not on how the code looks. Improve in place when the data model still fits the product, the defects are local (a screen, a rule, a message) and fixes stay fixed. Rebuild when a foundation is wrong in a way every feature inherits: no tenant key in a product that now has many customers, a data model shaped around a workflow you abandoned, or regressions that keep landing on the same invariants. Between the two sits a staged rebuild that replaces one workflow at a time behind the running app. Score the signals in the worksheet, and before any rebuild, list what the current app does that nobody wrote down, because a rebuild drops it unless someone carries it over.
Scope: the actor is the owner of an AI-generated app that has users or is close to launch, weighing whether to keep changing it or start again. The starting state is a working app with problems that are piling up. The boundary is the decision and the evidence behind it; running the rebuild is a separate project. The outcome is a written decision (improve, staged rebuild or full rebuild) with the three to five signals that settled it and a date to revisit.
Put it into practice
1. Sort recent problems into surface and foundation
List the last ten problems: bugs, changes that went badly, things users complained about. Tag each as surface (one screen, one rule, one message) or foundation (the data model, the permissions model, tenancy, how state is stored). Surface problems are what improving is for. A steady run of foundation problems is the rebuild signal.
2. Check whether the data model still describes the product
Products change shape: an internal tool becomes a customer-facing SaaS, a single-team app becomes multi-team. The tables may still describe the old product. A model that needs a tenant key it never had, or a many-to-many where it assumed one-to-one, touches nearly every query, and that is a foundation change however it's delivered.
3. Measure whether fixes stay fixed
From your change history or monthly release reviews, count changes that broke something that used to work, and note which invariant each broke. Regressions scattered across unrelated areas are ordinary maintenance. Repeated regressions in one area mean that area's structure is fighting every change made to it.
4. Test how expensive a small change has become
Make one small, tightly scoped change and look at what came back. If a one-behaviour request touches a dozen files or changes things nobody asked for, each improvement now carries review and re-test cost, and that cost belongs in the decision.
5. Inventory what the app does that nobody wrote down
Before any rebuild: the edge case someone handled months ago, the report a colleague runs on Fridays, the email that goes out when a record expires. Ask the people who use it daily, and run a parallel trial when replacing it. A rebuild recreates the documented app; the undocumented one is lost unless someone lists it.
6. Consider a staged rebuild
Rebuild one workflow at a time behind the running app and move users over workflow by workflow, instead of one big switch. It costs more coordination, and it lets you stop halfway if the new foundation turns out no better. The hard part to plan is data both versions can read while they run side by side.
7. Write the decision down with a revisit date
Improve, staged rebuild or full rebuild; the three to five signals that decided it; and when you'll look again (for improve, the next monthly release review). A written decision with its evidence is harder to reverse on a bad day than a feeling is.
8. Worked example (illustrative, synthetic data)
A fictional AI-built scheduling app made for one clinic, now being sold to others. Last ten problems: six surface (copy, a date picker, an export column), four foundation, all four tracing to appointments and patients having no clinic id. Schema review: no tenant key on five of seven tables. Regressions: three last month, all in appointment queries. A small change to the booking form came back touching 14 files. Undocumented behaviour: the front desk runs a weekly no-show report from a filtered view. Decision: staged rebuild, starting with the data model and the appointments workflow, while the original app keeps serving the first clinic until the new appointments workflow passes a parallel trial. Revisit in 30 days.
Improve-or-rebuild worksheet
Copy this structure into your review document and record your observed result for each row.
| Signal | Evidence to collect | Points toward | Worked example (synthetic) |
|---|---|---|---|
| Where recent problems live | Last ten problems tagged surface or foundation | Rebuild if foundation problems dominate | 4 of 10 foundation, all tenancy |
| Data model fit | Schema review against today's product | Rebuild if a key or relationship is missing across most tables | No tenant key on 5 of 7 tables |
| Regressions | Changes that broke working features, grouped by area | Rebuild the area if they cluster there | 3 last month, all appointment queries |
| Cost of a small change | What a one-behaviour request actually changed | Rebuild if small requests routinely come back large | 14 files for one form change |
| Permissions model | Two-account test results by role | Improve if failures are isolated gaps; rebuild if there's no consistent model | Passing for the one clinic |
| Code ownership | Code export acceptance test | Either; a failed export limits a developer-led rebuild | Export builds and runs |
| Undocumented behaviour | Interviews with daily users; a parallel trial | Improve if it is large and poorly understood | Weekly no-show report from a filtered view |
| Users and data at stake | Active users, record count, live integrations | Staged rebuild if real users depend on it daily | One clinic, two years of appointments |
| Decision | The option and the three to five signals that settled it | Improve, staged rebuild or full rebuild | Staged rebuild; revisit in 30 days |
A failure worth checking
The rebuild chosen because the code looked messy. The owner opens the export, finds long files and inconsistent naming, and decides to start again with a better prompt. The data model was sound and the real problems were four surface bugs. The rebuild takes weeks, reproduces the documented features without two undocumented behaviours users relied on, and ships something no better than before. Messy code that works on a sound data model is an improve decision. The counterexample: improving long past the point of sense. If every change to appointments breaks another appointment feature, a fifth careful fix is not cheaper than replacing the foundation it keeps colliding with.
Common questions
Isn't regenerating from a better prompt cheap with an AI builder?
Generating is fast; replacing is not. Most of a rebuild's cost is moving real data and users, rediscovering undocumented behaviour and re-verifying every requirement, and none of that gets faster because the first draft did.
Can I rebuild just the data model?
Sometimes. If the problem is a missing tenant key, adding it through a staged migration and updating the queries can be an improve decision with a large first step. It becomes a rebuild when most queries and screens have to change anyway.
Who should make the call?
Whoever owns the outcome, with the worksheet in front of them. If a developer or agency would do the rebuild, ask them to fill in the worksheet separately and compare; a disagreement on one signal is worth settling before committing.
Basis and scope
This is a proposed implementation method using illustrative examples, not a measured benchmark or a customer case study. Prepared with AI assistance. Validate product-specific behavior against current documentation and your own test environment.