The 25-Point AI App Health Check
A self-scoring audit for apps built with Lovable, Bolt, Replit, Base44 or Cursor. No developer required. Budget an afternoon.

This is the checklist we run internally before we quote on any AI app rescue job. It's been trimmed to the things a non-technical founder can verify without reading code, and ordered so the highest-consequence items come first.
Key takeaways
25 yes/no checks across 7 areas: data access, secrets, authorization, money, scale, abuse limits and operations.
Score 1 for yes, 0 for no or unsure. Most AI-built apps land between 11 and 16.
Any zero in Section 1 (data access) overrides your total — fix it first.
Everything here is fixable without a rewrite. The order of the list is the order to fix things in.
How to use it: work top to bottom, answer each honestly, and score 1 for yes, 0 for no or unsure. Unsure counts as no — that's not pedantry, it's the point. Anything you can't confirm is a thing you're trusting rather than knowing.
You'll need: your database dashboard, your live site, a browser, and a second test account.
Before you start: create that second test account now and log in with it in a private window. Half the checks below use it, and it's the single most revealing thing you'll do all afternoon.
Section 1 — Data access (5 points)
The highest-consequence section. If you only do one, do this one.
1. Row-level security is enabled on every table holding user data.
Open your database dashboard and look at the table list. Not "most tables." Every table with a user's anything in it.
2. Your access policies compare the row's owner to the current user.
Read the actual policy text. A policy whose condition is true, or that only checks the user is logged in, grants everyone access to everything. It looks enabled and does nothing.
3. A second account cannot see the first account's records.
Create a test user. Log in as them. Walk the app. If you see anything belonging to your main account, stop and fix this before anything else on this list.
4. Changing an ID in a URL doesn't return someone else's data.
Find a URL with a record ID in it. Change the number. Do it while logged in as the test account. If data loads, your boundary is cosmetic.
5. Deleted records are actually gone or actually retained — and you know which.
Soft deletes that still return in queries are a common quiet leak. Delete something as your test user, then check whether it still appears anywhere.
Why this fails so often: AI builders create tables fast and policies loosely, and the app's interface hides the problem — the data is readable, it just isn't shown. Most apps on Lovable and Bolt run on Supabase, where this is the single most common finding. More in The 6 Security Holes in Almost Every AI-Generated App.Scored below 5 here? See Supabase rescue or, if you suspect data has already leaked, data exposure fixes.
Section 2 — Secrets and keys (4 points)
6. No secret key appears in your page source.
Load your live site, view source, search for service_role, sk_live, sk-, SECRET, PRIVATE. Check the loaded JavaScript files too, not just the HTML. (In Chrome: DevTools → Sources → search across all files with Ctrl/Cmd + Shift + F.)
7. Your model provider key is server-side only.
If your app calls OpenAI, Anthropic or similar, that key must never reach the browser. Exposed model keys are found by scrapers within days and the first symptom is the bill.
8. Environment variables use the right prefix.
Anything prefixed for client use (NEXT_PUBLIC_, VITE_) gets inlined into the bundle you ship. Confirm nothing sensitive carries one.
9. Any key that was ever exposed has been rotated, not just deleted.
Removing a key from your code doesn't unpublish it. If it shipped once, it's public forever until you rotate it.
Found a key in the browser? Rotate it today, before anything else — then have the call moved server-side. Our security audit for AI-built apps covers secrets, policies and endpoints in one pass.
Section 3 — Authorization (3 points)
10. Admin routes reject non-admin users when typed directly.
Log in as your test user. Type your admin URL into the address bar. A page that renders at all — even broken, even empty — is a failure here.
11. Admin API endpoints check permission independently.
Harder to self-test. The proxy question: did anyone ever explicitly build server-side role checks, or was the admin link simply hidden from non-admins? If it's the latter, score 0.
12. Roles are enforced on the server, not stored in browser state.
If your app reads a role value into the frontend and decides permissions from it, that value is editable by the user holding it.
Logins also flaky? Authentication and authorization usually break together. See authentication fixes for AI-built apps, or Supabase auth not working for a single issue.
Section 4 — Money (4 points)
Skip this section entirely if you don't take payments — you'll score out of 21, so subtract 4 from every scoring band below.
13. Your webhook handler is idempotent.
Payment providers retry on timeout. If the same event arriving twice provisions twice or charges twice, you have a duplicate-charge bug waiting for a slow day.
14. A payment that succeeds while the user closes the tab still lands.
Test it: start a checkout, complete payment, close the browser before redirect. Then check whether your app knows.
15. A failed renewal revokes access.
Use your provider's test tools to fail a recurring charge. If the customer keeps full access afterwards, you're giving the product away to churned users indefinitely.
16. There's a path to issue a refund and reverse access.
Not an automated flow necessarily — just a defined way to do it that doesn't involve editing the database by hand.
Why this section matters more than its 4 points: checkout is the part of a payment flow everyone tests. Renewals, failures and retries are the parts nobody prompts for — and they fail silently, costing revenue for months before anyone notices. See payments rescue or Stripe webhook not working.
Section 5 — Scale and data volume (3 points)
17. You've tested at ten times your realistic first-year volume.
Not a million. Ten times what you honestly expect in twelve months. Seed the data and click through.
18. No screen loads every record and filters in the browser.
Symptom: a list that's instant at fifty rows and takes several seconds at a thousand.
19. You have a database migration path.
Can you change your schema without risking existing production data? If schema changes today mean editing and redeploying, score 0 — this is the condition that later makes every change frightening.
Outgrowing the builder itself? If scale problems trace back to the platform rather than your code, read Lovable vs Bolt vs Replit for Something You'll Actually Charge For, then see migrating off an AI builder.
Section 6 — Abuse and limits (3 points)
20. Login attempts are rate limited.
Try a wrong password fifteen times. If attempt sixteen behaves exactly like attempt one, there's no limiting.
21. Email-sending endpoints are throttled.
Signup, password reset, contact form. Unthrottled, these become a way to send thousands of emails through your provider and get your domain flagged.
22. File uploads have type and size limits, and stored files aren't publicly enumerable.
Try a very large file. Try a text file renamed to .jpg. Then open an uploaded file's URL in a private window — if it loads, your storage is public.
Want this checked properly? Self-tests catch the obvious cases. Vulnerability scanning and assessmentcatches the rest.
Section 7 — Operations (3 points)
The section almost nobody scores on, and the one that decides how bad a bad day gets.
23. Errors are captured somewhere you can read them.
Not shown to the user and lost. Captured, with enough context to reconstruct what happened.
24. You'd know about downtime without a customer telling you.
Alerting, uptime monitoring, anything. If the honest answer is "they'd email me," score 0.
25. You can roll back a bad deploy and restore from a backup you've tested.
Both halves required. A backup you've never restored from is a belief, not a safeguard.
Works in preview, breaks in production? That's usually an operations gap, not a code bug — see works in preview, breaks after deploy. For ongoing monitoring and fixes after launch, see application maintenance and support.
Your scorecard
Tally as you go:
Section | Max | Your score |
1. Data access | 5 | |
2. Secrets and keys | 4 | |
3. Authorization | 3 | |
4. Money (skip if no payments) | 4 | |
5. Scale and data volume | 3 | |
6. Abuse and limits | 3 | |
7. Operations | 3 | |
Total | 25 |
Scoring
22–25 — Production ready. Genuinely unusual for an AI-built app. Fix the stragglers and ship.
17–21 — Close. Usually a week or two of focused work. Nothing structural. Knock out the lowest-numbered gaps first.
11–16 — Typical. This is where most AI-built apps land. You have a working product and no safety net. Serviceable for a soft launch to friendly users, not for taking payments from strangers.
Below 11 — Don't take real customers yet. Not a judgement on the app. It means the app hasn't yet been given the parts that make failure survivable, and adding them now is dramatically cheaper than adding them after an incident.
Any zero in Section 1 overrides your total. A score of 20 with open database policies is a score of 20 with your customers' data readable by anyone who changes a number in a URL. Fix that first regardless of what the rest says.
What to do with your score
If you scored 17 or above, work down the list yourself. Every item here is fixable without a rewrite, and most are a few hours each. The order in this checklist is the order to fix them in.
If you scored below 17, or you hit several "unsure" answers in Sections 1 and 7, the useful next step is having someone read the actual codebase — not rebuild it, read it. Most of what matters shows up in an afternoon, and knowing where you stand is considerably cheaper than finding out from a customer.
If you've been trying to fix the gaps by prompting and each fix breaks something else, that's a sign of its own — see When to Stop Prompting and Start Engineering.
Where most apps lose points, by builder
Built with | Usually loses points in | Start here |
Lovable | Sections 1 and 3 (Supabase policies, auth) | |
Bolt | Sections 1 and 7 (unclaimed database, deploy config) | |
Replit | Section 7 (config living outside the repo) | |
Cursor / Claude Code | Sections 5 and 7 (no tests, no migration path) |
Frequently asked questions
Can I run this checklist without any coding knowledge?
Yes. Every check uses your database dashboard, your live site, a browser and a second test account. Items 11 and 19 are the hardest to self-verify; if you're unsure, score 0.
How long does it take?
About an afternoon for most apps. Section 1 alone takes under an hour and catches the highest-consequence problems.
Does a low score mean I need to rebuild my app?
Rarely. Almost every item is fixable in place, usually in hours rather than weeks. Rebuilds are only warranted when the data model itself is wrong — and even then, usually only part of the app.
Which section should I fix first?
Section 1, always. Then work through the list in order — it's sorted by consequence.
Is this checklist specific to one AI builder?
No. It applies to apps built with Lovable, Bolt, Replit, Base44, Cursor, v0, Claude Code or any similar tool. The gaps are the same because they come from what nobody prompted for, not from the tool.
Keep reading
Scored below 17, or answered "unsure" a lot?
We fill in this checklist against your actual code, with the specific files behind each answer. No obligation to hire anyone.
Built from findings across the AI-generated apps we've audited. Updated October 2026.



Comments