AI-generated app: 15 security checks before go-live

15 Controlli Sicurezza App Generata AI Prima Go-Live

AI-generated app: 15 security checks before go-live

When a web app, SaaS MVP, or internal tool created with AI seems to work, the natural pressure is to publish it. The problem is that “it works” and “it is ready for go-live” are two different conditions. Before connecting real data, users, payments, corporate APIs, or internal systems, you need a concrete verification of code, configurations, permissions, and exposed behavior.

This checklist is designed for teams that have used ChatGPT, GitHub Copilot, Cursor, Codex, Claude Code, Lovable, Bolt.new, Replit Agent, Devin, Kiro, Gemini, or other AI coding tools. Before go-live, what matters is what has been built, what is exposed, and what risks might arrive in production along with the feature.

Use it before go-live, before a client beta, before importing real data, or before connecting the app to payments, CRMs, corporate databases, cloud storage, or internal APIs.

The value of the checklist increases when it is applied to a real environment, not just the repository. Many vulnerabilities are not visible when reading a single function: they emerge when code, roles, routes, cloud configurations, storage, callbacks, and environment variables work together. For this reason, every check should produce verifiable evidence — an HTTP request, a configuration, a policy, a negative test, a screenshot of permissions, or a remediation decision.

🔴 Web Application Penetration Testing: identify hidden risks and strengthen your security with a focused assessment by ISGroup specialists.

Before starting: define the real perimeter

A checklist only works if the perimeter is clear. Before the 15 checks, gather at least this information: repository or project, staging and production URLs, public routes, APIs, user roles, processed data, authentication providers, databases, storage, cloud services, environment variables, dependencies added by AI, and parts generated or modified with assistants or agents.

If you cannot reconstruct these elements, the first risk is the lack of control over the go-live. An AI-generated app may have been built quickly, but before publishing it, it must become a verifiable product.

The perimeter must also include what was added to make the prototype work quickly: installed packages, temporary integrations, service accounts, test webhooks, buckets created during development, keys shared in chat, prompts used to generate code, and parts copied from examples. These are elements that often do not appear in the user story but can determine the app’s actual level of exposure.

What becomes public at go-live

The first check is to inventory what will be reachable: domains, subdomains, APIs, webhooks, callbacks, storage, preview deployments, admin panels, debug routes, static files, and temporary environments. Many incidents stem from what the team did not think they had published.

Check that demos, test endpoints, seed data, temporary admin accounts, diagnostic pages, and verbose error messages are not reachable. If you use Vercel, Replit, Lovable, Bolt.new, Firebase, Supabase, or similar clouds, also verify preview URLs, previous deployments, and custom domains.

This inventory must be compared with the product intent. If a route is public only because the framework exposed it by default, if a preview environment remains indexed, if a serverless function accepts calls without authentication, or if an admin endpoint responds outside the VPN, the attack surface is wider than expected. Before go-live, what is not needed must disappear; what is needed must have access control, logging, and consistent limits.

Authentication, sessions, and account recovery

Authentication must be verified beyond the “happy path” login. Test registration, login, logout, password reset, invitations, email changes, password changes, session expiration, and token revocation. If the app handles sensitive data, administrative roles, or payments, evaluate MFA or equivalent controls.

The most frequent errors in AI-generated apps are sessions that are too long, reusable password resets, tokens not invalidated after a password change, expired invitations that are still valid, checks only on the frontend, and error messages that reveal whether an account exists. Before publishing, every account flow must be tested with a valid user, an invalid user, an expired token, a link already used, and repeated attempts.

Also, check the behavior across devices and parallel sessions. A password change should invalidate tokens where necessary, a logout should truly close the session, and an invitation should not be usable by an account other than the intended one. In B2B applications, also verify what happens when a user leaves an organization, changes roles, or is removed from a tenant.

Authorizations, BOLA, and tenant isolation

Broken access control is one of the most critical risks for web apps and APIs. In apps created with AI, it often appears like this: the user is authenticated, but they can access objects that do not belong to them by changing an ID in the request.

Manually test user_id, project_id, tenant_id, organization_id, order_id, document_id, and every identifier controllable by the client. A standard user must try to read, modify, delete, or export records of other users; a user of one tenant must try to access resources of another tenant.

The check must be server-side. Hiding a button in the UI does not protect the API: every endpoint must verify identity, role, ownership, and tenant before returning or modifying data. AI-generated apps can introduce this problem in subtle ways — a new page uses a client-side filter, an export function forgets the tenant_id, a detail endpoint only checks if the user is logged in, a query uses an ID received from the browser without verifying ownership. The test must cross different flows, not just the main screen.

APIs and routes not visible in the UI

Many AI-generated apps have routes created for debugging, testing, or convenience that, even if they do not appear in the interface, can remain reachable via browser, curl, proxy, or scripts.

Call all APIs directly and try different HTTP methods, missing tokens, expired tokens, insufficient roles, incomplete payloads, manipulated parameters, and out-of-sequence requests. Verify admin routes, export endpoints, uploads, imports, webhooks, callbacks, health checks, and server-side functions. If a route should not exist in production, remove it; if it must exist, protect it with authentication, authorization, rate limiting, logging, and consistent error handling.

Also include in the check routes created by framework conventions: automatically generated API endpoints, dynamic pages, serverless functions, OAuth callbacks, payment webhooks, preview URLs, and cron job routes. A clean UI can hide a backend much larger than the team remembers.

Input validation, output handling, and file upload

Code generated by AI may correctly handle the intended path and completely ignore malicious input. Verify forms, query strings, JSON bodies, headers, filenames, paths, templates, markdown, HTML, and numerical values.

Minimum cases to test include SQL injection, NoSQL injection, XSS, path traversal, command injection, SSRF where relevant, active file uploads, double extensions, falsified MIME types, files that are too large, and malformed data. Validation must be server-side, with allowlists where possible. Output must also be managed: errors, logs, user messages, exports, and HTML rendering must not expose stack traces, tokens, queries, internal paths, or other users’ data.

For uploads, verify the entire file lifecycle: loading, scanning or validation, saving, downloading, previewing, deleting, and accessing via shared links. If the app generates previews, PDFs, images, or documents, check that the uploaded content is not interpreted as code or reused in prompts, templates, or automated workflows without sanitization.

Secrets, API keys, and credentials

Search for secrets in code, .env files, Git history, prompts, chats, logs, build outputs, artifacts, tests, READMEs, examples, and configurations. AI-generated apps often contain keys inserted “temporarily” to make a demo work.

If a key has been exposed in a repository, prompt, log, or artifact, it must be rotated: removing it from the file is not enough. Use secret managers or managed environment variables, separate dev/staging/prod, and assign minimal scopes. Beware of the frontend: a private key must not end up in the bundle, on a page, in a JSON response, or in a client-side log.

Verification must include the final build and artifacts, not just source code. Search for keys in the minified bundle, source maps, pipeline logs, containers, exported configuration files, and examples left in the repository. If an AI assistant suggested pasting a key to “test quickly,” treat it as potentially compromised until it has been rotated.

Personal data, corporate data, and environments

Before go-live, you must know what data you process, where it ends up, and who can read it. Personal data, customer documents, application logs, webhook payloads, tickets, screenshots, exports, and backups must have a clear location.

Demos generated with AI often use real data too soon. Separating development, staging, and production reduces risk: use synthetic data when possible, avoid real dumps in repositories, and verify log retention. If the app handles personal data, check the legal basis, privacy policy, minimization, access, deletion, retention, and involved vendors. This does not replace a privacy assessment, but it avoids publishing without knowing where data flows.

A useful check is to follow a single piece of data from insertion to deletion: forms, APIs, databases, logs, emails, notifications, webhooks, analytics, backups, and exports. This path quickly shows if the app retains more information than necessary or if production data ends up in development environments, support tools, tracking systems, or prompts used for debugging.

Database, BaaS, and security rules

If you use Supabase, Firebase, Replit Database, managed SQL databases, or backend-as-a-service, it is not enough for the app to read and write correctly: you must verify access rules.

  • For Supabase: check Row Level Security, policies by role and tenant, service role keys never in the client, SQL functions, storage policies, and anon keys.
  • For Firebase: check Firestore/Realtime Database Rules, Storage Rules, custom claims, and default deny.
  • For SQL databases: check queries, user_id/tenant_id filters, migrations, backups, and network access.

The most important test remains manual: two different users, two different tenants, similar objects, direct APIs, and attempts to read or modify unauthorized data. Do not assume that BaaS rules are equivalent to application logic: a page may show only the correct data, but a permissive policy may allow direct access from the client. Conversely, an overly rigid policy may push the team to use server-side service keys without granular controls. Before publication, consistency between rules, code, and the data model is required.

Storage, buckets, and uploaded documents

Uploads and storage are often added in a hurry by AI app builders. Verify who can upload, read, download, delete, and share files, because a public bucket or a predictable URL can expose corporate documents even if the rest of the app requires login.

Check the separation between public assets and private files, MIME types, maximum size, extensions, path traversal, antivirus or equivalent checks when necessary, link expiration, backups, and deletion. Try downloading and deleting with different users: one user must not be able to download another customer’s files just by modifying IDs, paths, or filenames.

Also verify metadata. Original filenames, authors, sizes, paths, tags, previews, and extracted text can reveal sensitive information even when the main file seems protected. If uploaded documents are processed by AI functions, check where extraction results are saved and who can read them.

Dependencies, lockfiles, and supply chain

AI tools often add packages to resolve errors quickly. Every change to package.json, requirements.txt, pyproject.toml, go.mod, Cargo.toml, lockfiles, and install/build/test scripts must be reviewed.

Check for known vulnerabilities, package maintenance, licenses, transitive dependencies, typosquatting, postinstall scripts, locked versions, and near-homonymous libraries. If a dependency was added only to make a test pass, ask if there is a simpler or already approved solution. SCA and dependency reviews help, but they do not replace technical decisions: does that library really need to enter the product?

When AI modifies dependencies, also look at what is removed. An upgrade can change security middleware, parsing, validation, cookie management, or default behavior; a library replaced with a simpler alternative may lose controls that the team took for granted. The review must include lockfiles and configuration diffs, not just application code.

Cloud, deploy, and configurations

Deployment is the point where shortcuts and defaults become real exposure. Verify CORS, CSP, security headers, callbacks, redirects, debug mode, error reporting, rate limits, environment variables, IAM, security groups, buckets, public databases, containers, serverless functions, and IaC.

An AI-generated configuration can resolve a build error and weaken security at the same time. CORS *, non-allowlisted redirects, active debug, detailed errors, secrets in deployment logs, and overly broad cloud permissions are problems to block before publication. If the project uses AWS, Google Cloud, Firebase, Supabase, Vercel, Replit, or other managed environments, also check dev/staging/prod separation and service account privileges.

Deployment configurations deserve a test in an environment similar to production: verify real headers, effective redirects, error behavior, published assets, API responses, allowed callbacks, and access from external networks. A correct configuration file in the repository does not guarantee that the platform is running the app with the same values.

CI/CD, branch protection, and artifacts

The pipeline must prevent code generated or modified by AI from reaching production without review. Check branch protection, mandatory reviewers, minimum tests, secret scanning, dependency scanning, build artifacts, CI/CD logs, and token permissions.

Beware of changes that “simplify” the pipeline: disabled tests, removed security steps, tokens with broader scopes, automatic deployment on unprotected branches, more verbose logs, and caches containing sensitive data. If you use agents that open PRs or modify pipelines, require human approval on sensitive areas: auth, APIs, secrets, dependencies, IaC, deployment, and data.

Also check what is kept as an artifact. Reports, builds, coverage, logs, test dumps, screenshots, and compressed packages can include data or configurations that should not leave the pipeline. A fast but overly verbose pipeline can turn into a second surface of exposure.

Negative tests and abuse cases

AI-generated tests often cover the happy path. Before go-live, you need tests that try to break the rules: users without permission, different tenants, expired tokens, malicious input, invalid uploads, race conditions, out-of-order flows, manipulated payments, reused invitations, and abused password resets.

Every critical check should have at least one positive and one negative test. If a feature involves data or roles, it is necessary to demonstrate not only that the right user succeeds, but that the wrong user fails. Manual WAPT is useful because it brings a different mindset: it does not ask if the feature works, but how it can be abused from the outside.

For every important flow, prepare at least one abuse question: can I create something without paying, see another user’s data, repeat an action, skip a step, force a state, use an old token, change price, modify quantities, bypass a limit, or export more data than expected? The answers to these questions matter more than the percentage of test coverage.

Logging, monitoring, and audit trail

Logs and monitoring are needed to understand what happens after go-live, but they can become a risk if they contain tokens, PII, sensitive payloads, queries, stack traces, or other users’ data.

Verify which critical events are recorded: login, password reset, role change, admin actions, export, delete, payment, upload, errors, configuration changes, and anomalous access. Logs must be useful for incident response, but minimized and protected. If the app is created with AI agents, also keep evidence of what was generated, which PRs were accepted, which tests passed, and which remediations were made: the audit trail helps reconstruct decisions and responsibilities.

Monitoring must be ready before launch, not added after the first incident. Define which events produce alerts, who receives them, which thresholds indicate abuse, and which actions are planned. For an exposed web app, failed logins, anomalous 403/401 errors, spikes on sensitive endpoints, unusual uploads, massive exports, and repeated password resets are signals to be treated with attention.

Remediation decision before go-live

The checklist is useless if it does not produce a decision at the end. Every finding must have a status: fix before go-live, accept temporarily with justification, monitor, or block publication.

Block the go-live if you find unauthorized access to data, exposed secrets, unexpected public buckets, reachable databases, APIs without auth, privilege escalation, manipulatable payments, or overly permissive cloud configurations.

You can plan for after launch only what has a clear residual risk, a defined owner, a deadline, and monitoring. If no one owns the risk, the risk is not accepted: it is just postponed. A good remediation decision distinguishes between functional defects, exploitable vulnerabilities, hardening, and technical debt: not everything blocks the release, but what concerns data, authorizations, secrets, payments, public surfaces, and privileges must have priority. Go-live should be based on closed findings or explicitly governed risks, not just the perception that the prototype works.

When an internal review is enough and when an independent verification is needed

An internal review may be enough for non-public prototypes, without real data, without roles, without exposed APIs, without payments, without corporate integrations, and with a reviewer competent in the modified areas.

Independent verification is needed when the app is online or about to be, handles real data, has external users, manages payments, uses administrative roles, integrates corporate APIs, exposes uploads, modifies cloud/IaC, or derives from large PRs generated by AI agents. The pressure of go-live must not disappear: it must be transformed into priority. First, check data, authorizations, secrets, exposed surfaces, and production configurations; then decide what can go online.

How ISGroup can verify an AI-generated app

The check changes based on what has been built and what is exposed. If the app or APIs are reachable online, Web Application Penetration Testing verifies real behavior from the outside. If the risk is in the code, application logic, middleware, secrets, or dependencies, Code Review helps identify vulnerabilities and regressions before merging. If the problem concerns known surfaces, hosts, services, and exposed configurations, Vulnerability Assessment helps map known vulnerabilities and priorities.

If the AI-generated app has… Main risk Recommended check
Web app, API, public routes, uploads, login or payment flows Behaviors abusable from the outside Web Application Penetration Testing
Generated code, auth, roles, business logic, secrets, dependencies Vulnerabilities or regressions in the code Code Review
Hosts, services, versions, exposed configurations, technical endpoints Known vulnerabilities or technical exposures Vulnerability Assessment
Cloud, IAM, buckets, databases, CI/CD, IaC, service accounts Misconfiguration or excessive privileges Cloud Security Assessment
Trust boundary, multi-tenant, integrations, sensitive data Weak architectural assumptions Secure Architecture Review
Continuous use of AI coding in the release cycle Non-repeatable checks on releases and pipelines Software Assurance Lifecycle

The choice of check depends on what has truly changed: code, exposed behavior, cloud, architecture, or development process. Before go-live, it is advisable to delimit that perimeter and verify the actual risk on the application.

Have you created an app with AI tools and need to connect it to real data, users, or payments? ISGroup can help you verify code, APIs, authorizations, secrets, dependencies, cloud, and exposed surfaces before publication.

Evidence to prepare before the review

Prepare the repository or project, environment URLs, list of parts generated with AI, user roles, APIs, auth providers, databases, storage, cloud services, added dependencies, environment variables, available logs, pipelines, remediation decisions, and already accepted risks. This evidence makes the verification faster and more concrete because it allows distinguishing application bugs from misconfigurations, code problems from cloud risks, and blocking findings from plannable improvements.

Frequently Asked Questions

  • Are automatic tests enough before go-live?
  • No. They are necessary, but they do not prove by themselves that authorizations, tenant isolation, business logic, APIs, and configurations are secure in real behavior.
  • Do I always need to do a WAPT?
  • If the app or APIs are exposed online, WAPT is the most direct check on external behavior. If the risk is in code not yet exposed, it may be more consistent to start with a Code Review.
  • Is a no-code or low-code app generated with AI more secure?
  • Not automatically. Even if the code is hidden or partially managed by the vendor, there remain configurations, auth, roles, databases, storage, API keys, callbacks, and real data to verify.
  • Which findings block go-live?
  • Unauthorized access to data, privilege escalation, exposed secrets, unexpected public databases or buckets, APIs without auth, manipulatable payments, and overly permissive cloud configurations.
  • When is a Cloud Security Assessment needed?
  • When the app uses cloud, BaaS, managed databases, storage, IAM, service accounts, CI/CD, IaC, or deployments configured quickly with the help of AI.

Protect your organisation with Web Application Penetration Testing.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert

Useful sources and references