Turn AI Prototype Into Production: What Breaks
How to turn an AI prototype into production: the five failures that hit first, and how to decide between hardening a vibe-coded app and rewriting it.
AI Engineer, Viithiisys

How do you turn an AI prototype into production? The five things that break, in order
Authorisation breaks first, then migrations, database queries, exposed secrets and error visibility. Each one passes a demo untouched and fails once real users, real data and a second deploy arrive.
To turn an AI prototype into production, fix these in the order they cause damage, not the order you notice them. Veracode's 2025 GenAI Code Security Report tested more than 100 LLMs and found 45% of generated code samples failed its security tests. Generated code is therefore production ready only after a person has reviewed it, whichever tool wrote it.
| Failure | When it shows | Blast radius | Fix type |
|---|---|---|---|
| Missing auth and permissions | First hostile or curious user | Data leak | Policy rewrite |
| No migrations or rollback | Second schema change | Outage, data loss | Tooling plus backfill |
| N+1 queries | Around a few hundred users | Slow pages, timeouts | Query changes, indexes |
| Secrets in the client bundle | Anyone opens dev tools | Account takeover, bills | Rotate, move server-side |
| No error visibility | First silent failure | Churn you never see | Monitoring setup |
Auth and permissions: the one that leaks data
Auth is the first failure because a prototype checks that a user is logged in, not what that user may touch. The result is a working login screen sitting over a database that answers any request.
What does missing row-level security look like?
Many AI builders, Lovable among them, wire the front end straight to a hosted Postgres through a public key. That is safe only when row-level security policies decide what each request can read. The Supabase documentation is explicit that tables exposed through the API need RLS enabled and policies written.
When the policies are absent or too broad, the damage is public. The National Vulnerability Database entry for CVE-2025-48757 describes insufficient row-level security in Lovable-generated projects that let unauthenticated attackers read or write arbitrary tables.
How do you test authorisation in an afternoon?
Create two ordinary accounts and a third with no role. Then log in as each and request the other's records by changing an ID in the URL, the API call or the browser console.
Repeat for every table and every write operation, not only reads. Test the admin screens as a non-admin, and test with the session token removed. Anything that answers is a finding, and every finding goes on a list before any feature work resumes.
No migrations, no rollback
Prototypes change the schema by clicking in a dashboard or letting the AI edit tables directly. Without versioned migrations, no one can recreate the database, review a change or undo it.
This stays invisible until the second environment. Staging does not match production, a column rename works on one and fails on the other, and nobody can say which change caused it.
The fix is to capture the current schema as a baseline migration, then put every later change in version control and run it through CI. Tools such as Prisma Migrate, Flyway and Alembic all do this. Pick the one that matches the stack and stop hand-editing production.
What does a rollback you can actually run look like?
A rollback is a tested path back, not a hope. For each migration, write the reverse step or an explicit note that it is irreversible, such as a dropped column.
Before any destructive change, take a backup and restore it into a scratch database to confirm it works. Many teams own backups they have never restored, and the first restore attempt should not happen during an incident. Add the destructive steps to a release checklist so that nobody runs them on a Friday evening.
Why do N+1 queries only show at 200 users?
An N+1 query runs one query for a list, then one more per row, so 50 items make 51 round trips. With 5 test rows it feels instant; at a few hundred users the database becomes the bottleneck.
With 5 test rows nobody notices the cost, because each extra round trip takes a few milliseconds. Multiply that by concurrent users and by the number of rows per page, and the same code that felt instant in a demo starts timing out under real traffic.
AI-generated code produces this pattern easily, because each component fetches its own data and nothing sees the whole page. A dashboard with a list of orders, each loading its customer, then its items, then its status, multiplies quickly.
The Rails guides on eager loading describe the same problem and the standard remedy: load associations in one batched query. The principle holds in any ORM or hand-written SQL.
How do you find them before your users do?
Turn on pg_stat_statements in Postgres and sort by total execution time and call count. A query called thousands of times per page view with trivial individual cost is the signature.
Then load-test the three busiest screens with a few hundred simulated users on production-sized data. Fix by batching or joining, add indexes on the foreign keys you filter by, and paginate every list. Re-run the same test to confirm the numbers moved.
Secrets in the client bundle
Anything shipped to the browser is public. If an API key sits in front-end code, anyone can read it from the network tab or the built JavaScript in under a minute.
Build tools make this easy to get wrong. The Vite documentation states that variables prefixed with VITE_ are exposed to client source code, and other frameworks use similar prefixes. A key placed behind such a prefix is published, not hidden.
Public keys are fine when designed for it, such as a Supabase anon key protected by RLS. The dangerous ones are service-role keys, payment secrets, LLM provider keys and third-party tokens. An exposed LLM key is the costly one, because strangers can spend against your account.
What is the order of operations when a key has leaked?
Rotate first, move second. Treat every key that ever touched the client as compromised, revoke it at the provider, and issue a new one.
Then move the call behind a server route or edge function that holds the secret and enforces its own limits per user. Search the git history as well as the current code, because deleted keys stay in old commits. Finally, add a secret scanner to CI so the same mistake cannot return unnoticed.
No error visibility, so you find out from customers
A prototype has console logs and nothing else. In production, a failed payment webhook or a crashed form submission leaves no trace, and the first report is an angry email.
Silent failure is also what makes the other four problems expensive. A leaking policy, a slow query and a broken migration all look identical from the outside: a user who quietly leaves.
What is the minimum observability worth having?
Start with four things:
- Error tracking with source maps, so stack traces point at readable code instead of minified bundles.
- Structured logs with a request ID, so one failed request can be followed across the front end, the API and the database.
- An uptime check on the health endpoint, so an outage is noticed in minutes rather than by the next customer.
- An alert that reaches a human, by phone or chat, not an inbox that nobody opens.
Add per-request timing so slow queries surface as numbers, not complaints. For AI features, log prompts, model, latency and token counts, with personal data redacted. Decide who is on the hook when the alert fires, because an alert with no owner is a dashboard nobody reads.
What does it cost to turn an AI prototype into production?
It costs a fixed block of senior engineering time, scoped to the size of the app and the five areas above. Viithiisys does not publish a price, because the scope depends on what the review finds.
A review is quoted as a fixed scope after a discovery call, so the number is agreed before work starts. The commercial shape matters more than a figure: a defined list of checks, a written findings report and a clear split between hardening work and rewrite candidates.
What does the review cover?
The review covers:
- Authorisation and data access
- The schema and its migration history
- Query behaviour under load
- Secret exposure and deploy configuration
- Logging and alerting
It also covers dependency and licence checks, backups and environment parity.
The cheapest time to find an authorisation hole is before the first customer, and the most expensive is after the first screenshot.
The output is a ranked list. Each finding gets a severity, a fix and a view on whether it can be patched in place. Teams that want to keep building meanwhile can bring in QA and testing or DevOps consulting for the parts they cannot staff.
Rewrite or harden: how do you decide?
Harden when the data model matches the business and the faults sit at the edges. Rewrite when the schema or the architecture works against the product you now need to build.
Viithiisys has shipped software since 2007, with more than 500 projects delivered for clients including Paytm, Snapdeal, IKEA, Nestle, Shiprocket and Vikram Solar. The pattern on rescue work is consistent: the decision turns on the data model and the coupling, not on how the code looks.
When is hardening the right call?
Hardening fits when tables map to real entities, the business logic is mostly in one place and the stack is common enough to hire for. Typical work is writing policies, adding migrations, batching queries, moving secrets and wiring monitoring.
The trade-off is that you inherit generated structure. Naming is inconsistent, tests are rare, and a few files may be very large. Budget time to add tests around the paths that handle money and identity before changing them.
When is a rewrite the right call?
Rewrite when the schema stores state in ways that cannot represent your real rules, such as one user per account when you need teams. The same applies when business logic is scattered across UI components, or the platform locks you into a runtime you cannot scale.
A rewrite keeps the prototype's real value: validated screens, user feedback and a clear spec. For a fresh build, the Moonship MVP programme is the route to start again against an agreed, fixed scope. Larger systems go through custom software development with engineering from Mohali and a Canadian office in Markham, Ontario.
What is the next step?
Before you commit to either path, get an outside engineer to read the code against the five areas above. A short review costs far less than discovering a leaked table through a customer's email.
If the prototype is already live, ask for the authorisation and secrets checks first, since those carry the legal and financial exposure. To get that written down with a recommendation, get a production-readiness review and bring the repository, the deployment details and a list of what real users do today.
FAQ
- How long does it take to turn an AI prototype into production?
- It depends on how much of the five failure areas is already sound, which only a code review shows. Hardening a small app is often a matter of weeks, while a rewrite takes longer. Viithiisys scopes this after a discovery call and does not publish a standard duration.
- Is AI generated code production ready out of the box?
- Not by default. Veracode's 2025 GenAI Code Security Report found 45% of AI-generated code samples failed its security tests. Generated code usually works on the happy path, so authorisation, migrations, query performance, secrets handling and monitoring all need human review before real users arrive.
- Can a Lovable app run in production?
- Yes, if it is hardened first. Lovable apps commonly use Supabase, so the key checks are row-level security on every table, no private keys in the client bundle, a migration history, and error monitoring. A publicly disclosed 2025 vulnerability involved missing row-level security in generated apps.
- Should I rewrite my AI-built prototype or harden it?
- Harden it when the data model is sound and the problems sit at the edges, such as policies, secrets and logging. Rewrite when the schema fights the real business rules or logic is tangled into UI components. A short code review settles which case applies.