Audit & fix
The 20 ways AI-built apps fail in production
Apps built with Lovable, Bolt, Cursor or Replit tend to fail in the same ways once real users arrive. The twenty failure modes we check for, and what fixes each one.
In short
AI coding tools are very good at the happy path: one user, one browser, everything working. Production is everything else, and AI-built apps fail there in the same ways again and again — open database access, secrets in the browser, payments that charge without updating anything, no backups, no monitoring, and store rejections for mobile. These are the twenty failure modes we check every AI-built app for, grouped by area, with what each one looks like and what fixes it.
Key takeaways
- Most serious failures are about who can reach which data, not about features.
- Payments usually “work” in a demo and break on refunds, retries and cancellations.
- An app without monitoring and a tested restore is one bad deploy from a long outage.
- Mobile apps built from web code fail store review for predictable, fixable reasons.
Why do AI-built apps fail in the same ways?
Because AI coding tools optimize for what you can see. A prompt produces a screen that works when you click through it, and that is a genuine achievement. What you can’t see from the screen — whether the database checks who is asking, whether a webhook is really from Stripe, what happens when a deploy fails at 2 a.m. — is exactly what the tool had no reason to get right.
That is also why the list is predictable. The same gaps appear whichever tool built the app, because they are the parts no demo exercises.
What are the twenty failure modes?
Grouped by area. Each one says what goes wrong, and then what fixes it.
Security
- Your database is publicly readable. Row-level security was never switched on, so the browser key that ships in your JavaScript bundle can read every row in every table — including other customers’ records. From the outside the app looks completely normal. What we do Access policies written and tested per table, with a test asserting a user of one tenant cannot reach another’s rows, run in CI so a future migration cannot silently remove isolation.
- A secret key is sitting in your JavaScript bundle. An admin-level key was used in client code to work around access rules that were never configured. It bypasses every policy you add later, so fixing the policies alone changes nothing. What we do Rotate the key, move every privileged call behind server-side endpoints, and add a build-time scan of the client bundle that fails the build if a key pattern reappears.
- Admin buttons are hidden, not blocked. The role check lives in the React component. The endpoints behind it have none, so anyone who calls them directly can perform administrative actions — including changing their own role. What we do The role check moves to the server on every privileged endpoint. The interface still hides what you cannot do, but hiding is no longer what stops you.
- One customer can read another customer’s data. The tenant identifier used to filter queries is read from a value the client sends. Change it in the request and another organization’s records come back, because the server never checks it against the session. What we do Tenant identity resolved server-side from the authenticated session only, and the client-supplied parameter rejected outright rather than ignored.
- Uploaded files are public to anyone with the URL. The storage bucket was left open and filenames are predictable. Invoices, identity documents and private images are readable without logging in, and enumerable by anyone who guesses the pattern. What we do Per-object access rules, signed URLs with an expiry, and non-guessable keys — plus a sweep of what is already exposed.
Payments
- Payments charge but nothing happens. The card is charged, the customer sees a success screen, and the account is never upgraded. Almost always the webhook is unverified, unhandled, or never arrives at a handler that does anything. What we do Signature verification on every event, the payment provider treated as the source of truth for subscription state, and a reconciliation pass over what is already wrong.
- One customer, three subscriptions. Payment providers retry webhooks by design and occasionally deliver duplicates in normal operation. The handler inserts a row each time, so your database quietly disagrees with the provider about what was bought. What we do Event IDs stored with a unique constraint, and subscription writes made upserts keyed on the provider’s own identifier rather than blind inserts.
- Cancellations and refunds never reach your database. One event is handled — the successful checkout — and nothing else. Canceled customers keep their access, refunded customers still count as revenue, and failed renewals are invisible. What we do The full lifecycle implemented and tested: renewal, failure, refund, dispute, cancellation, and the grace periods between them.
- The price is set by the browser. The amount to charge is sent from the client to the checkout session. Anyone can change it in the request before it is sent, and pay whatever they like. What we do Price resolved server-side from your own catalog. The client sends an identifier for what it wants to buy, never a number.
Reliability
- Works locally, fails in production. Different environment variables, a missing build step, a server-only library imported into client code, a database connection that works from your laptop and not from the host. The error is a 500 with nothing behind it. What we do Environment parity, configuration validated at boot so a missing value fails loudly at deploy rather than silently at runtime, and real errors surfaced instead of swallowed.
- You find out it is down from a customer. No error tracking, no uptime check, no alerting. A broken payment path or a failing migration is reported by whoever happens to hit it, and there is no trace left to diagnose from afterwards. What we do Error tracking with releases and source maps, uptime checks on the paths that matter, and alerts routed somewhere a human actually reads.
- It got slow and nobody knows why. Missing indexes, queries issued inside a loop, whole tables fetched and then filtered in memory. All of it is fine with fifty rows and none of it is fine with fifty thousand. What we do The slow queries identified from real timings, indexes added where they are needed, repeated queries collapsed, and a performance check that fails the build on a regression.
- There is no way back from a bad deploy. Deploys go out from a laptop, there is no record of what shipped, and the only way to undo a release is to write the old behavior again by hand under pressure. What we do Deploys from CI on merge, immutable builds you can identify, and a rollback that has been executed at least once so you know it works.
Data
- No backups, or none anyone has restored. Running on defaults with no export schedule, or with backups nobody has ever restored from. A backup that has not been restored is an assumption, not a recovery plan. What we do Scheduled backups with a retention policy, a documented restore, and that restore actually performed into a scratch environment to prove it works.
- Staging writes to the live database. One connection string across every environment. Preview deployments mutate real customer data, and a migration tested on a branch runs against production. There is also nowhere safe to rehearse a restore. What we do Separate projects per environment with distinct credentials, and a seeded non-production dataset containing no real customer records.
- Schema changes are made by hand. Columns added in a dashboard, no migration files, no history. Environments drift apart and nobody can recreate the database from scratch — including whoever has to recover it. What we do Migrations committed to the repository and applied in CI, so the schema is reproducible from zero and every change is reviewable.
Mobile
- Rejected by the App Store for minimum functionality. A web view with nothing a browser could not do. Resubmitting the same build with a better description does not change the outcome, and each attempt costs a review cycle. What we do Genuine native capability — push, biometric unlock, offline behavior, native share — implemented against the real platform APIs, which is what the guideline actually asks for.
- Refused at upload for an old target SDK. Google raises the target API floor every year and refuses new builds below it before review even begins. The upgrade then surfaces layout and orientation changes the app never handled. What we do Target level raised and the consequences fixed: edge-to-edge layout with correct insets, and large-screen orientation behavior that no longer assumes portrait.
- Your privacy answers do not match what the app sends. The data safety form declares no collection while the app ships analytics and account data. That mismatch is itself a policy violation, independent of the collection it fails to declare. What we do Declarations rewritten from an audit of the network calls the build actually makes, plus privacy manifests and required-reason declarations for every bundled dependency.
- Push and deep links work in the simulator only. Certificates, entitlements and associated domains were never configured for a release build, so notifications silently fail and links open the browser instead of the app. What we do Configured and verified on physical devices against a release build, on both platforms, with the release process written down so you can repeat it.
Twenty is what every audit checks for, not an exhaustive list. Every app also has failures specific to its own stack and history — finding those is what an audit is actually for.
Which should you fix first?
- Anything that exposes data. Open database tables, secret keys in the bundle, one customer reading another’s records. These are the ones that end up in an incident report.
- Anything that loses money. Unverified webhooks, a price set by the browser, subscriptions that never reach your database.
- Anything that makes recovery impossible. No backups, no restore test, no way back from a bad deploy.
- Then reliability and store readiness, in the order your launch needs them.
How can you check your own app?
Start with the first two areas above, using the production checklist for Lovable apps — most of it applies to any AI-built app. If you’d rather have it done for you, the AI-Built App Audit Report covers all twenty with read-only access, ranks every finding by severity with evidence and a fix estimate, and is delivered in two business days.