Skip to main content
TACAVAR
•Build in Public

Vibe-Coded App Production Hardening Checklist

What to audit before an AI-built prototype takes real users: secrets, auth, LLM I/O, evals, cost loops, retention, and the observability to ship.

**Short answer:** Before an AI-built prototype takes real users, audit seven things: production control and ownership, secrets, auth and data-layer access, LLM input/output validation, eval coverage, runaway-cost loops, and data retention — then add the minimum observability that tells you when something breaks before a customer does. Audit per module, not pass/fail on the whole app. If the data model can't hold what the product needs next, no amount of patching the other six saves it.

This is the checklist we run at Tacavar on every inherited codebase, in the order we run it. The order matters: each step is the safety net for the one after it.

Why inheritance is a different problem

Vibe coding got easy in 2025. Taking over what it produced did not. By the end of 2025, roughly 41% of all code written globally was AI-generated, and GitHub Copilot alone writes close to 46% of the average developer's code. The consequence shows up at handover: of roughly 10,000 startups that shipped AI-built apps this way over the past year, more than 8,000 now need partial rebuilds or rescue engineering at $50,000 to $500,000 each, according to two independent 2026 analyses. Security researchers at Escape.tech scanned 5,600 live, publicly deployed vibe-coded apps and found over 2,000 high-impact vulnerabilities and 400 exposed secrets — roughly one in three shipped with a serious, exploitable flaw, most often missing access control or an unvalidated webhook.

The mechanism is structural, not sloppiness. Each prompt that generated the app optimized for what was in front of it. Nothing enforced a plan across prompts. The parts of the system that never appear in any single prompt — permission logic behind the data layer, webhook signatures, what happens when two users hit the same record at once — are exactly the parts that are missing. And the cost compounds: the same analyses estimate every month spent adding features on an unaudited foundation adds 20–30% to the eventual rebuild bill, because each feature deepens its dependency on the structure that will need untangling.

"It demos" and "it can be maintained safely" are two different claims. The checklist below converts the first into evidence about the second.

Step 0: Map control before you touch code

The first audit isn't code. It's ownership. Fast-built apps routinely end up fragmented: repo in one personal account, cloud under a different email, domain and DNS with a founder's old registrar, database and payments under someone who left. Inventory who holds the repo, cloud, DNS, database, environment variables, third-party APIs, payments, email, monitoring, and backups — and whether you can log in to each and transfer it. Do not redeploy anything before this map exists. A one-line change you can't roll back is a risk, and you can't know it's rollable until step 1 confirms it.

Step 1: Reproducible build and rollback

Can you rebuild and deploy from source in a clean environment? Can you return to the last working version when a deploy fails? If the deploy process lives on one laptop or in one person's memory, every subsequent change is a bet. This is the baseline requirement in the AWS Well-Architected Operational Excellence pillar — repeatable, reversible change — and it's the safety net for everything after it. Until it works, freeze feature work.

Step 2: Secrets audit

Scan the repo and the frontend bundle for hardcoded API keys, database passwords, and provider tokens. Fast builds hardcode credentials in code, config, or — worst case — client-side JavaScript. If they're in git history or shipped in a public bundle, treat them as exposed regardless of whether you've seen abuse: rotate first, then move everything into environment variables or a secrets manager. OWASP classifies use of hard-coded credentials as a software weakness, and its Top 10 covers the adjacent security-misconfiguration risk. Rotation order: anything reachable from the client bundle first, database credentials second, everything else after.

Step 3: Auth and data-layer access

This is where the Escape.tech finding concentrates — missing access control was the most common serious flaw. Check every route for an auth gate, not just the ones with a login screen. Then check the layer below: does data access enforce tenant scoping, or does one logged-in user's query touch another tenant's rows? Vibe-coded apps frequently authenticate the route and skip authorization at the data layer, because that's the part no prompt could see. Add webhook signature verification while you're here — an unvalidated webhook is an unauthenticated write endpoint wearing a different URL.

Audit per module. The right pattern assigns each module one of three verdicts: keep (checks pass), fix-in-place (specific failing checks, architecture stays), or rebuild. The branch that matters most: if the data model itself can't support the product's next phase, the verdict is rebuild — no amount of patching auth checks changes that. Everything else is a days-long fix.

Step 4: LLM input/output validation

If the app calls a model, the model is now part of your attack and failure surface. Three checks:

  • **Input:** can user-controlled text reach the prompt unchecked? That's a prompt-injection surface — someone can redirect the model's behavior through content. You don't need a full defense architecture on day one; you need to know where the surface is and not give it write access to anything expensive.
  • **Output:** is model output validated before it becomes a database write, an API call, or rendered HTML? Unvalidated LLM output is unvalidated user input with extra steps — the model can be made to emit any string.
  • **Evals:** is there any golden set at all — a handful of known inputs with known-good outputs you can re-run after every prompt or model change? Without one, you can't tell whether the model quietly regressed. Prompt behavior will drift the first time someone upgrades the model or rewrites the prompt from memory.

If your agents also act — touching tools, spend, or infrastructure — that's a governance problem on top of an inheritance problem. We cover the control layer for that separately in our agent firewall teardown; it decides which tools an agent can touch, what it can spend, and what it can write. This checklist gets an inherited app safe; the firewall keeps it safe once autonomous agents are operating inside it.

Step 5: Cost loops and data retention

Two silent failure modes live here.

**Runaway cost loops.** Agent and LLM code paths fail in a specific way: a retry loop, an unbounded tool call, or an agent that re-runs a task on every webhook delivery. Nothing crashes; the invoice just grows. Before shipping, put a cap on every model call path — max iterations, max spend per task, max daily budget with a hard stop. Qualitative floor: if you cannot answer "what is the worst possible bill this code can produce in a day?", you haven't finished this step.

**Data retention.** Vibe-coded apps rarely decide what they keep. Audit what user data, model inputs, and model outputs are stored, where, and for how long. Prompts frequently contain things users wouldn't put in a form. Decide retention deliberately — forever-by-default is a liability and a compliance problem you inherit along with the code.

Step 6: The minimum observability to ship

Not a monitoring platform. The minimum: error monitoring with alerts landing somewhere a human sees them, a log line per model call and per external API call, and a kill switch — a flag that turns off the expensive or risky path without a deploy. Before any refactoring, add the layer that means "someone will know when it breaks." Every change before that is a blind change, and users notice failures before teams do.

This is the dividing line between an inheritance audit and how you *run* AI systems day to day. We run our own operations through coordinated Claude Code sessions with defined topology and human checkpoints — orchestration, not audit. Don't build the second before the first exists.

The order, compressed

| Step | Audit | Verdict question | |------|-------|------------------| | 0 | Control & ownership | Can you log in to everything and transfer it? | | 1 | Build & rollback | Can you deploy from source and roll back? | | 2 | Secrets | What's hardcoded, exposed in history, rotated? | | 3 | Auth & data access | Is every route gated and every query tenant-scoped? | | 4 | LLM I/O & evals | Is model input/output validated; is there a golden set? | | 5 | Cost loops & retention | What's the worst daily bill; what's retained and why? | | 6 | Observability | Does a human know when it breaks, without a customer? |

When the map is done, the decision is usually one of four: maintain directly, partial refactor, gradual replacement, or full rebuild. Not every inherited app needs a rewrite — but you can only see which one is justified after the inventory, never before.

Frequently Asked Questions

Should I rebuild a vibe-coded app from scratch or harden it?

Decide per module after the audit, not by reputation. If the data model can't hold the product's next phase, that module (or the app) is a rebuild candidate. Everything else — auth gates, webhook signatures, secret rotation, error handling — is fix-in-place measured in days.

When should I pause shipping new features?

Until steps 0–2 (control, rollback, secrets) are done. Every feature added on an unaudited foundation increases the eventual rebuild cost — the 2026 rescue analyses put the drag at an estimated 20–30% per month of feature work on top of unaudited structure.

How is this different from an agent firewall or MCP security checklist?

This is a one-time inheritance audit for a codebase going to production: ownership, secrets, auth, LLM I/O, cost caps, retention, observability. An agent firewall is the standing control layer for agents that act on tools and spend after they're live. An MCP tool-call security checklist is narrower still — per-call authentication, scoping, and logging for the Model Context Protocol surface. They stack; they don't substitute.

What does a hardening audit cost in time?

The audit itself is a week-scale exercise: the control map and secrets scan are days, the per-module auth and LLM I/O checks are the bulk of it. The expensive part is the verdicts — but per the rescue analyses above, every month of feature work shipped before the audit adds an estimated 20–30% to the eventual rebuild bill, so the audit is the cheap end of the decision. If you're choosing between orchestration frameworks or rebuild strategies afterward, our agent-framework comparison covers the landscape.

How Tacavar Growth helps

We run this audit on every inherited codebase before touching the code — and we've inherited enough of them to know the pattern: the two modules that will corrupt customer data in month three are findable in week two, if you look per module instead of trusting the demo. If your team needs governed content and operations systems built on that discipline rather than bolted on after, start with Growth Starter — $497/mo, self-serve, or request a strategy scan. You built it. We optimize it.

Related Reading

AI Agent Infrastructure: Prototype to ProductionRunning a Company on Coordinated Claude Code SessionsAI Crypto Trading Bot Fees: What You Actually Pay