You described a website to a model and got a working website. Not a mockup — a real one, with a database behind it, a login, a payment form, and a URL you can send to people.
That is a genuine achievement and this post is not going to pretend otherwise. You compressed weeks of work into an afternoon and you learned whether anyone wants the thing, which is the most expensive question in software to answer.
But there is a specific reason a site like that starts generating problems around the time it starts getting real visitors, and it is not that the code is bad. It is that the code answers the question you asked, and production asks different questions. Nobody prompted for “and make sure a stranger cannot read another customer’s invoices”, because that is not a feature. It is a property. Properties do not come from prompts.
Here is the whole shape of the problem before the detail.
Why this happens, in one chart
The single most useful piece of evidence here is Veracode’s ongoing measurement: they generate code from a large pool of models against fixed tasks, then scan the result. In the Spring 2026 update — 150-plus models, 80 tasks, four languages — the security pass rate was 55%, which they describe as virtually identical to two years earlier.
The interesting part is not the average. It is the spread.
Look at what models are good at and what they are not. Injection and weak cryptography are near-solved: the fix is short, canonical, and appears thousands of times in training data. Cross-site scripting and log injection collapse to 13-15%, because whether a value needs escaping depends on where it came from and where it is going — context that exists in your application and not in the prompt.
That is the exact signature of a vibe-coded codebase. The famous vulnerabilities are handled. The situational ones are not.
Two other numbers frame this honestly, and both come from people who use these tools daily rather than from vendors selling against them:
- In the Stack Overflow 2025 Developer Survey (49,000+ respondents), 84% use or plan to use AI tools — and the number-one frustration, cited by 66%, is solutions that are “almost right, but not quite”. 45.2% say debugging AI-generated code takes longer than writing it themselves would have.
- DORA’s 2025 report, from nearly 5,000 professionals, found AI adoption correlates with higher delivery throughput and lower delivery stability, at the same time. Their framing is the one worth keeping: AI doesn’t fix a team; it amplifies what’s already there.
So the tools are fast, the output is plausible, and plausible-but-wrong is the expensive failure mode. None of that argues for not using them. It argues for knowing which parts nobody has checked.
Gap 1: the trust boundary is in the wrong place
This is the big one. If you read nothing else here, read this section.
When a model writes “only paying customers can export data”, it satisfies that instruction the most direct way available: a conditional in the interface. The export button disappears for everyone else. Demo it and it works perfectly.
Everything on the left of that dashed line runs on a computer the visitor owns. They can read it, edit it, and skip it. A paywall in JavaScript is a paywall that asks politely.
The version of this that ends in a news story is not clever. Someone opens the network tab, watches which request the app makes, and replays that request without your interface in front of it. If the check lived in the interface, there is now no check.
This is not theoretical, and the scale is documented. Escape scanned publicly reachable vibe-coded applications and reported 2,038 highly critical vulnerabilities and 400+ leaked secrets across 1.4K applications, including 175 separate instances of exposed personal data, bank account data among them.
A note on that figure: several write-ups of this research say 5,600 apps were scanned. Escape’s own page says the findings are across 1.4K applications, so that is the number used here. Where a secondary source inflates a primary one, take the primary.
What “fixing it” actually means
Not adding a second check in the UI. Restating every rule where the data lives:
- Supabase or Firebase: row-level policies that deny by default, so an unauthenticated read returns nothing instead of the table. If you have ever seen a tutorial say “disable RLS for now to get it working”, that is the line that has to be undone.
- Your own API: every endpoint independently establishes who is calling and whether they may do this, per request. Not once at login.
- The service key — the one that bypasses all of it — never reaches the browser under any circumstances, including inside an environment variable prefixed for client use.
Gap 2: the keys are in the page
Related, but a separate fix, and usually more urgent because it is already public.
Nothing in that chain is an attack. Bundlers inline client-side environment
variables because that is what they are for — NEXT_PUBLIC_ and VITE_ prefixes
are documented promises to ship the value to the browser. A model asked to call
Stripe or OpenAI from a component will put the key where the call is.
The part owners consistently underestimate is the right-hand side of that figure. Rotating a leaked key takes minutes. Answering “what was accessed while it was live” requires request logs from before you knew there was a problem — and a prototype does not have them. That is the question a regulator, an insurer or an enterprise customer will ask, and “we rotated the key” is not an answer to it.
Fix: move every third-party call behind your own server endpoint, rotate the keys afterwards, and check the provider’s usage dashboard for the exposure window while you still can.
Gap 3: one environment, no backups
In July 2025 an AI agent on Replit ran destructive commands against a live production database during an explicit code freeze, then reported that a rollback was impossible. The platform shipped automatic development/production separation afterwards — it had none at the time.
Treat that as the reference case rather than the cautionary tale, because the condition it reveals is nearly universal in prototypes: one database, serving both the thing you are experimenting on and the thing customers depend on.
Three things, cheap, in this order: a separate production database that development cannot reach; automated backups; and one restore you have actually performed. An untested backup is a belief, not a backup.
Gap 4: the site cannot be told apart from itself
Now the commercial gap, and the one we see most often as a paid rebuild.
If your site renders in the browser, every URL may arrive at Google as the same empty shell with the same title, the same description and the same canonical.
Google does execute JavaScript — it renders in a queue with headless Chromium, so
this is not a claim that client-rendered sites cannot rank. It is a narrower and
more damaging claim: pages that are not distinguishable cannot rank separately.
Google’s own guidance is explicit that routes behind a #fragment are not
reliably parsed as separate URLs, that you should not use JavaScript to change a
canonical after the fact, and that SPA error routes need either a real redirect or
an injected noindex to avoid registering as soft 404s.
We know this one from the inside. The previous version of this site was a React single-page app where every service sat behind a fragment on one route. Google canonicalises fragments to the root, so the entire company website was four indexable URLs. No amount of content would have fixed that; the architecture was the ceiling. Rebuilding it is what this site now is.
Gap 5: fast on your laptop
Core Web Vitals are assessed on real visits, at the 75th percentile, split between mobile and desktop.
That percentile is the part that surprises people. It is not the average visitor — it is roughly the slowest quarter of them who decides. A generated app that loads a large JavaScript bundle to render text is comfortable on a developer machine on wired broadband and poor on a three-year-old Android phone on mobile data.
The same pass picks up accessibility, and in Ontario that is not optional. Under
the AODA, since 1 January 2021, designated public sector organisations and
businesses or non-profits with 50 or more employees must meet WCAG 2.0
Level AA on public websites, with narrow exceptions for live captions and
pre-recorded audio descriptions. Generated markup routinely fails the easy parts
of this: unlabelled inputs, plain div elements used as buttons, contrast that looks fine and
measures 3:1, focus states removed for aesthetics.
Gap 6: why the next change costs more than the last one
This is the gap owners feel without being able to name it. Version one took a weekend. Version four takes three weeks and breaks two things.
GitClear classified 211 million changed lines from 2020 to 2024. Copy/pasted lines went from 8.3% to 12.3% of all changes, while moved lines — refactoring, the fingerprint of someone noticing two things are the same and making them one thing — fell from about 25% to under 10%. In 2024, copy/pasted lines exceeded moved lines for the first time in the dataset.
Generated code has a structural reason to do this. A model asked for a feature writes that feature. It does not know that a near-identical function already exists three files away, so it writes a second one. Do that forty times and your password rule lives in six places — and the day you change it, you change four of them.
The order to fix it
Most audits hand you a list sorted by how interesting each item is. This one is sorted by what cannot be undone.
The order to fix a vibe-coded website, from irreversible risks to long-term maintainability
- 1Revoke and re-home every secretStop the bleedingHours
A key in a shipped bundle is already public. Rotate it, move the call server-side, and check the provider’s usage logs for the period it was exposed.
- 2Put authorisation behind the APIStop the bleedingDays
Every rule that currently lives in a component has to be re-stated where the data is. With Supabase or Firebase that means row-level policies that deny by default.
- 3Separate environments and take backupsStop the bleedingHours to days
One database serving development and production is the single condition that turns a mistake into an outage. Add a restore you have actually tested.
- 4Turn on error and uptime monitoringBefore more users arriveHours
Until something reports failures, your detection system is a customer’s email. This is the cheapest item on the list and it changes every item after it.
- 5Give every page its own URL, title and canonicalBefore more users arriveDays to weeks
Architectural, not cosmetic: pages that cannot be told apart cannot rank separately, and no amount of content fixes it.
- 6Fix the field-data vitals and accessibilityBefore you spend on trafficDays
Measured on real visits, not a local audit. Keyboard access and contrast come with this pass, and in Ontario they may be a legal requirement.
- 7Make it changeableSo the next change is cheapOngoing
Tests around the paths that move money or data, one deploy pipeline, and consolidation of the duplicated logic. This is what makes the next six months cheap.
Data that leaks stays leaked. A slow page is fixed next week with no residue. So secrets, authorisation and backups come before performance, and performance comes before elegance — regardless of which one is more fun to work on.
What a professional team actually changes
Not “rewrite it”. Rewriting is usually the wrong call: your prototype encodes real decisions about what the product is, and those were expensive to learn. What changes is the part that was never in the prompt.
A boundary you can point at. Security stops being scattered checks and becomes one place where authorisation is decided, so a new feature inherits it instead of re-implementing it.
Enforcement instead of intention. This is the habit that transfers best from our own work. On this site, a page description outside 70-165 characters fails the build. A missing canonical fails the build. A diagram without a text alternative fails the build. Those are not code review comments someone can merge past — they are gates. Conventions written in a document decay; conventions written in a script do not.
Environments, monitoring and a tested restore, so the first person to know about a problem is you and not a customer.
An architecture that matches the job. Interactive product? Keep the app framework. Content and marketing pages that need to be found? Those should be pre-rendered, one URL per intent, and that is usually the largest single performance and visibility win available.
Consolidation with tests around the money paths, so the sixth change costs what the first one did.
What to do this week, free
Before hiring anyone, including us:
- Open your site, then view source. Search the page and the JS bundle for
key,secret,sk_andtoken. If you find something live, rotate it today. - Open the network tab, find a request that returns data belonging to a logged-in user, and replay it signed out. If data comes back, that is your first priority.
- In a private window, check whether three different pages have three different
titles. Then search Google for
site:yourdomain.comand count the results. - Run the page through PageSpeed Insights and read the field data section, not the lab score.
- Try the whole site with the keyboard only — tab, enter, escape. If you cannot reach a button, neither can a real subset of your customers.
Five findings from twenty minutes tells you whether this is a weekend of hardening or a rebuild, and you will know more about your own product either way.
If you would rather have someone go through it properly, that is the web development work we do — and if the problem turns out to be that nobody can find the site, that is SEO and Google ranking. Building a mobile app the same way has a different set of gates, most of them enforced by two stores: that is the companion post.