Production readiness for AI-built apps

Your Claude app works.
Now make it survive real users.

AI gets you to a demo in a weekend. Getting from there to something customers can trust takes judgment about auth, security, SEO, deployment, and failure. I'm Noah Chait, and that's the senior-engineer work I do.

readiness-report.logauditing
  • Authorization enforced server-sideFAIL
  • Secrets out of the client bundleFAIL
  • Prerendered HTML + per-route metadataWARN
  • CI gates every mergeFAIL
  • Backups restore-testedWARN
  • Errors page a humanFAIL
  • LLM spend capped per userWARN
  • One-command rollbackFAIL

Readiness

14/100

Illustrative: the shape of an audit, not a real client.

01The gap

A model optimizes for the demo path. Production is every other path.

What AI-generated code is great at

  • Getting a feature working end to end, fast
  • Plausible UI, plausible API, plausible schema
  • The happy path, with one user, on localhost

What it can't see from inside a prompt

  • Who shouldn't be allowed to call this endpoint
  • What happens at 3am when a dependency is down
  • How anyone finds, deploys, monitors, or rolls it back

I use these tools daily. The missing pieces are the ones nobody notices until a customer, a crawler, or an attacker does, and a senior or staff engineer owns that layer.

02Common pitfalls

Eight places AI-built apps quietly fall over.

Most apps have a few of these, and a good audit mostly orders them by what would hurt most.

Auth & authorization

Who are you, and what may you touch?

01

How it fails

  • Login works, but permissions are only enforced by hiding buttons in the UI
  • Sequential IDs let any signed-in user read someone else’s records
  • Tokens in localStorage, no expiry, rotation, or revocation story

What I put in place

  • Server-side authorization on every route, tested with cross-tenant attempts
  • Hardened sessions (httpOnly, rotation, revocation) or a managed IdP with SSO
  • Roles and tenancy modeled in the data layer, not bolted on in the front end

Security

Assume the internet is hostile.

02

How it fails

  • API keys (including LLM keys) shipped in the client bundle or committed to git
  • No input validation, wide-open CORS, no rate limiting, no security headers
  • Untrusted text flowing into prompts and tools with no injection boundary

What I put in place

  • Secrets in a manager, keys proxied server-side, history scrubbed and rotated
  • Schema validation at every boundary, CSP and header baseline, per-user rate limits
  • Threat model for the agent surface: least-privilege tools, output handling, audit logs

SEO & discoverability

If crawlers see an empty <div>, so does Google.

03

How it fails

  • Client-rendered SPA: the HTML that crawlers and link previews fetch is blank
  • One title and description for every route; no canonical, sitemap, or robots
  • Slow first paint and layout shift tanking Core Web Vitals

What I put in place

  • SSR, SSG, or prerendering matched to how often the content changes
  • Per-route metadata, canonical URLs, Open Graph, structured data, sitemap
  • Performance budget: image pipeline, font loading, caching and CDN headers

CI/CD & environments

Deploys should be boring.

04

How it fails

  • Production is whatever was last pushed from someone’s laptop
  • No tests gating merges, no preview environments, no way to roll back
  • Config drift between local, staging, and prod; migrations run by hand

What I put in place

  • Pipeline that lints, type-checks, tests, and builds on every pull request
  • Preview deploys, environment parity, migrations as a reviewed, gated step
  • Progressive rollout with a rollback you have actually rehearsed

Data & persistence

The schema outlives the prototype.

05

How it fails

  • Schema shaped by whatever the last prompt produced; no migration history
  • Backups that were never restored, or never configured at all
  • Missing indexes and N+1 queries that only show up at ten users

What I put in place

  • Versioned migrations, constraints, and a data model you can reason about
  • Automated backups with a restore drill and defined recovery targets
  • Query review, indexing, and connection pooling sized to real load

Observability & incidents

You can’t fix what you can’t see.

06

How it fails

  • console.log is the monitoring stack
  • Users find outages before the team does
  • No runbook, no on-call, no idea what “normal” looks like

What I put in place

  • Structured logs, tracing, and error tracking wired to real alerts
  • A few SLOs that matter, dashboards a human can read at 3am
  • Runbooks for the failures you’ve actually seen, plus a blameless review habit

LLM reliability & cost

The model is a dependency like any other.

07

How it fails

  • Unbounded tokens, no timeouts or retries, no per-user spend cap
  • Prompts changed by feel; no evals, so regressions ship silently
  • Hard-coded model IDs that will be deprecated on someone else’s schedule

What I put in place

  • Budgets, caching, backoff, and graceful degradation when the API is down
  • An eval set that gates prompt and model changes like any other test suite
  • A thin provider boundary so upgrading a model is a config change

Testing & code health

So the next change doesn’t break the last one.

08

How it fails

  • Zero tests; correctness depends on whoever is re-reading the diff
  • 3,000-line files and three copies of the same logic
  • Dependencies pinned to whatever the model remembered

What I put in place

  • Characterization tests around what works, then refactor behind them
  • Clear module boundaries, lint and type gates, dependency update automation
  • Docs and ADRs so the team can own the code without me

Also on the list when it matters: privacy and compliance, accessibility, cost management, dependency upkeep.

03Self-audit

Two minutes. Twelve honest questions.

Check only what's true today. Nothing leaves your browser.

Auth
Security
SEO
CI/CD
Data
Ops
LLM

04What a senior engineer adds

The value is judgment about the code you already have.

01

Order the risks

Hardening means deciding what to fix first, what to accept, and what can wait. You get a ranked list, not a 200-item scan.

02

Right-size the answer

A 50-user internal tool doesn’t need Kubernetes. A multi-tenant SaaS needs more than a shared admin password. I match the solution to the stakes.

03

Write the decision down

Short ADRs and runbooks so the “why” survives. Future engineers, and future AI prompts, inherit the context.

04

Leave your team stronger

Pairing, code review, and handoff are part of the work. The goal is a codebase your team can run and extend without me.

05How we'd work

Audit, harden, hand off.

  1. 01

    Readiness audit

    I read the code, the infrastructure, and the deploy path. You get a written, risk-ranked report with each fix sized, so you can decide what to do with or without me.

  2. 02

    Harden

    I fix the top risks in your repo through reviewable pull requests: auth, security, CI/CD, SEO, observability, without a parallel rewrite.

  3. 03

    Hand off

    Runbooks, docs, and a walkthrough with your team. Optional light-touch advisory afterward for the questions that come up in the first months.

Scope and pricing depend on the size of the app, so I quote after a short intro call. You can start with the audit alone.

Ready when you are. Is your app?

Tell me what you built, what it runs on, and what worries you. I'll reply with an honest read on where it stands.