Fix a Vibe Coded Application: 22 Checks
How to fix a vibe coded application for production: a 22-point checklist covering security, tests, dependencies, licences and deploys, plus keep-or-replace rules.
AI Engineer, Viithiisys

How do you fix a vibe coded application: rewrite or repair?
Fix it first. A vibe-coded app often has a validated product idea with weak engineering underneath, though that varies by project. To fix a vibe coded application, audit it against a fixed checklist, keep what passes, and rewrite only the parts that fail structurally.
Why does the rewrite instinct usually fail?
A rewrite throws away the one thing the prototype proved: that users want the behaviour. It also resets your bug count to unknown, and the new code is often written with the same tools and the same blind spots.
The evidence on those blind spots is consistent. Pearce et al. tested GitHub Copilot on 89 security-relevant scenarios in 2021 and found roughly 40% of the 1,689 generated programs were vulnerable (arXiv:2108.09293).
A Stanford study by Perry et al. in 2022 found participants using an AI assistant wrote less secure code and were more confident it was secure (arXiv:2211.03622). Taken together, the two studies suggest that confidence without verification is a recurring risk in vibe coding.
When is a rewrite justified?
Rewrite when the foundation is wrong, not when the surface is ugly. Three conditions qualify: a data model that cannot represent your real entities, authentication bolted on at the UI layer only, or a framework choice you cannot hire for.
Everything else is refactoring. Duplicated components, inconsistent naming and 600-line files are annoying but fixable module by module. The checklist below separates the two cases so the decision rests on a score, not on how the code feels to read.
What is on the 22-point checklist to fix a vibe coded application?
Twenty-two checks across six areas: security, data, code structure, tests, dependencies and operations. Each has a pass condition you can verify in under an hour, so the audit produces a score rather than an opinion.
The 22 checks, with pass conditions
Security
- Authorisation enforced server-side. Pass: every endpoint rejects another user's record IDs.
- Secrets out of the repo and bundle. Pass: no keys in git history or client JavaScript.
- Input validation and parameterised queries. Pass: no string-built SQL, and schema validation on every API input.
- Rate limiting on auth and costly endpoints. Pass: login, reset and LLM-calling routes are throttled.
- Session and token handling. Pass: expiry, rotation and revocation work.
- Row-level or tenant isolation. Pass: tenant A cannot read tenant B in a test.
Data
- Schema under migrations. Pass: changes are applied by versioned migration files only.
- Backups with a tested restore. Pass: a restore has been run, with a recorded time.
- Constraints and indexes. Pass: foreign keys, uniqueness rules and indexes on queried columns.
Code structure
- Single source of truth for business rules. Pass: pricing, permissions and status logic live in one place.
- Error handling and logging. Pass: no swallowed exceptions, and structured logs carry request IDs.
- No dead or duplicated code. Pass: unused routes, files and copied components are removed.
- Typed boundaries. Pass: API contracts are typed or schema-checked.
Tests
- End-to-end test on the revenue path. Pass: signup to payment passes in CI.
- Tests on integrations. Pass: payment, email and webhook handlers are covered, including failures.
- Regression test per fixed bug. Pass: each fix ships with a failing-then-passing test.
Dependencies
- Vulnerability scan clean or triaged. Pass: no unreviewed critical advisories.
- Licences reviewed. Pass: no unapproved copyleft in shipped code.
Operations
- Reproducible build and deploy. Pass: one command or merge deploys, with no manual steps.
- Separate staging and production. Pass: different databases, keys and domains.
- Monitoring and alerting. Pass: errors, latency and uptime alert a human.
- Rollback path. Pass: the previous release can be restored in minutes.
How should you score it?
Score each check as pass, partial or fail, and weight security and data failures as blockers. Any fail in checks 1, 2, 6 or 8 means the app should not take new users until fixed.
The NIST Secure Software Development Framework, SP 800-218, is a useful reference for the underlying practices, because it describes outcomes rather than tools (NIST SSDF). Treat the checklist as a practical subset of it, sized for a small product rather than a regulated enterprise.
Which code can you keep and which must you replace?
Keep code that is correct, isolated and testable. Replace code that mixes concerns, hides business rules in the UI, or sits on a data model that cannot be migrated. The decision is per module, not per application.
What signals code worth keeping?
Look for pure functions, clear inputs and outputs, and logic that already matches the business rule when you read it against a spec. UI components that only render props are cheap to keep even when they are verbose.
Generated CRUD screens, form layouts and marketing pages usually fall here. They are low-risk, rarely contain security decisions, and can be tidied gradually. A practical test: if you can write a test for the module in ten minutes without mocking half the app, keep it.
What signals code to replace?
Replace modules where authorisation, pricing or state transitions are scattered across components and handlers. These are the places where AI ai-generated code cleanup turns into archaeology, because the same rule exists in four slightly different versions.
Also replace anything holding secrets client-side, any hand-rolled authentication, and any database schema without migrations or constraints. For structural cases like these, custom software development with a defined architecture is faster than patching, because each patch adds another divergent copy of the rule.
How do you refactor AI-written code safely?
To refactor AI written code without breaking it, pin behaviour first. Write characterisation tests that record what the module does today, including its bugs, then restructure while those tests stay green.
Change behaviour in a separate commit from structure. Mixing them is how a tidy-up quietly alters a pricing calculation, and nobody notices until an invoice is wrong.
How do you add test coverage from zero without stopping shipping?
Start with one end-to-end test on the path that earns money, run it in CI, then add a test with every bug fix. Do not chase a coverage percentage; chase coverage of the failures that would cost you customers.
Which tests come first?
Order them by blast radius. First, a browser-level test of signup through payment. Second, tests on webhook handlers and third-party integrations, including the failure branches. Third, authorisation tests proving one user cannot read another's data.
Those three catch the failures that end up in support tickets and incident reviews. Unit tests on utility functions feel productive and protect very little.
If your own developers wrote or prompted the code, they share its blind spots, so they tend to test the paths they already believe work. An independent QA and testing team can write the regression suite and CI wiring from the spec rather than from the code, while your developers keep shipping features.
How do you keep releases moving meanwhile?
Run the new suite in report-only mode for a week, fix flakiness, then make it a required check on the main branch. Block merges only on the tests you trust.
Adopt a rule for new work: every bug fix and every touched module gets a test. Coverage then grows where the code is actually changing, which is where regressions happen, and nobody has to pause the roadmap for a testing sprint.
How do you audit dependencies and licences?
Run an automated vulnerability scan, then review licences manually for anything you ship or distribute. Generated projects tend to pull in many packages, some unmaintained, and a few the code never uses.
How do you handle vulnerable and abandoned packages?
Start with the built-in tooling. The npm CLI documents npm audit for checking installed packages against known advisories (npm audit docs). GitHub documents Dependabot for automated alerts and update pull requests (Dependabot docs).
Triage matters more than the raw count. An advisory in a build-time tool is not the same as one in your request handler. Remove unused packages first, since every deletion shrinks the attack surface for free.
What licence risks should you look for?
Generated code can introduce packages under copyleft terms such as the AGPL, which may impose obligations on a hosted product. Use a licence scanner for your ecosystem to produce a full inventory, then have someone decide the policy.
Also check for pasted snippets with licence headers or attribution comments, and for any model-output provenance you cannot explain. If the product is heading into acquisition or enterprise procurement, expect a buyer to ask for exactly this inventory, so produce it now.
What does the infra and deploy pipeline need?
A repeatable pipeline: build from a clean checkout, run tests, deploy to staging, promote to production, and roll back in minutes. Anything that depends on one person's laptop fails this check.
What is the minimum pipeline?
Source control with protected main, CI that runs lint, type checks and tests, infrastructure defined as code, and secrets in a managed store. Separate staging and production databases and keys.
The AWS Well-Architected Framework sets out operational excellence and reliability practices in this shape, and it is a reasonable yardstick even if you do not host on AWS (AWS Well-Architected). For teams without in-house pipeline experience, DevOps consulting can stand it up in a short, bounded engagement.
What do you monitor?
Error rate, latency, uptime and, for any LLM-backed feature, token spend per user. Alert a person, not a dashboard nobody opens.
Two failure modes are worth alerting on specifically: a runaway loop that burns API budget overnight, and a database that fills its disk because nothing prunes logs. Both are cheap to detect and expensive to discover from an invoice.
When should you hand it to a real engineering team?
Hand it over when the app has paying users, handles personal or payment data, or the person who prompted it can no longer explain how it works. At that point the risk is operational, not creative.
What should you prepare?
Repository access, the prompt history or specs if they exist, a list of integrations with credentials rotated, and a plain statement of what the app must do. Include the three or four workflows that make money.
A good team will ask for read-only access first and return a scored audit before proposing work.
How does Viithiisys approach a rescue?
Viithiisys works from the same checklist shown above. It starts with read-only access, scores the 22 checks, and shares the written result before any fix work is quoted. Blockers in security and data come first, and rebuild decisions are made per module from the score rather than from a general preference for starting over.
What if the core needs rebuilding?
Where the audit says the foundation fails, rebuild the core against a fixed scope rather than patching indefinitely. Moonship is Viithiisys's fixed-scope MVP build service. Scope and timeline are agreed before work starts, which suits a rebuild of a validated prototype.
Where the issue is system-level, such as a monolith nobody can change, system modernisation is the more fitting route. The honest trade-off: a rebuild costs calendar time that a patch does not, and it only pays back if the audit found structural faults.
What does this cost and how long does it take?
Cost and duration depend on app size, data model quality, integration count and existing test coverage. Viithiisys scopes each engagement after a discovery call and quotes a fixed scope afterwards, so no figure is published here.
What drives the scope?
Four things move it most: how many user roles and tenants exist, how many third-party integrations carry money or personal data, how much of the data model needs migration, and whether anything is tested today.
A single-user internal tool with one database is a small job. A multi-tenant SaaS with payments, file uploads and an LLM feature is a different category. The audit exists to tell you which one you have before you commit budget.
How is the engagement structured?
In three steps: a written audit scored against the checklist, a prioritised fix plan that puts blockers first, and then delivery against a defined scope. The audit is the cheap step that prevents the expensive mistake, whether that is an unnecessary rewrite or a missed security hole.
If your own team prefers to run the checklist first, do it. If you want an independent read on what is actually broken, book a broken workflow assessment and bring the repository.
FAQ
- How do I fix a vibe coded application without rewriting it?
- Audit first, then fix in order of risk. Close security and data-loss gaps, add end-to-end tests around the revenue paths, move deploys into CI, and refactor modules only when you next change them. A full rewrite is justified only when the data model or auth design is unsound.
- Is AI-generated code safe to run in production?
- Not by default. Peer-reviewed studies, including Pearce et al. (2021) and Perry et al. (2022), found AI-assisted code frequently contained vulnerabilities. It can ship safely after review, dependency scanning, secrets handling and tests, the same gates any untrusted contribution would face.
- How long does an AI generated code cleanup take?
- It depends on size, data model quality and how much is already tested. The audit comes first and produces a scored report. Timelines and fixed scope are then set per project after a discovery call, because a prototype with one integration and a multi-tenant SaaS differ enormously.
- Should I hire a freelancer or an engineering team to fix my app?
- A freelancer suits a bounded fix, such as one broken integration. A team suits an app with paying users, because security review, QA, infrastructure and ongoing ownership are separate skills. Ask for a written audit and a defined scope before any work starts.